About the Role

We are building ID1X — a large-scale identity intelligence platform that ingests, deduplicates, normalizes, and fuses data from dozens of OSINT vendors into a unified identity graph. The data pipeline processes terabytes of structured and unstructured data monthly through a Medallion architecture (Bronze → Silver → Gold), feeding downstream AI agents, a knowledge graph (Neo4j), a vector database (PGVector), and a unified search engine (OpenSearch/M-Mind).

This is a senior individual contributor role. You will architect and build the backend services and data pipelines that power the entire platform — and you will mentor junior engineers on engineering standards, data modeling, and system design. The role has meaningful architectural influence from day one.

Key Responsibilities

  • Design and implement Dagster-orchestrated data pipelines across the full Medallion stack: Bronze (raw ingestion) → Silver (normalization, dedup, enrichment) → Gold (identity graph population, vector embedding generation)
  • Build and own the vendor data middleware layer: normalize OSINT vendor outputs (District4, IntelliX, Bright Data, Social Links, Constella) to a canonical internal schema with mandatory fields, source quality scores, and confidence weighting
  • Implement deduplication and identity resolution logic: Redis-backed fingerprinting for idempotent ingestion, Confluent Kafka for high-throughput streaming at 1TB+/month inbound
  • Design and optimize PostgreSQL schemas for job state management, identity tracking, and audit logging; implement read replicas, partitioning, and migration strategies
  • Build and maintain the ACP (API Control Plane): query planner, normalization layer, adapter pattern for provider-agnostic data access, Swagger documentation, P99 ≤ 20ms performance baseline
  • Integrate with Neo4j knowledge graph: design entity/relationship schemas, Cypher query optimization, graph enrichment pipeline
  • Integrate with PGVector for semantic embedding storage and similarity search; coordinate with AI/ML engineers on embedding generation models
  • Implement OpenTelemetry trace propagation (Trace ID / Flow ID / UUID) end-to-end across all backend services — critical for support and incident debugging
  • Enforce backend engineering standards: code review, SonarQube quality gates, test coverage, CI/CD discipline via GitHub Actions
  • Mentor junior and mid-level engineers; contribute to architecture decision records (ADRs) and system design documentation

Tech Stack You'll Own

  • Language Python (primary) — expert level. FastAPI. Pydantic v2. asyncio. Type annotations throughout.
  • Pipeline Dagster (orchestration). Medallion architecture (Bronze/Silver/Gold). Provider-agnostic ingestion design.
  • Messaging Confluent Kafka (streaming, 1TB+/month). RabbitMQ (job queuing). Redis (dedup fingerprinting, caching).
  • Storage PostgreSQL (HA, partitioning, read replicas). MinIO (object storage).
  • OpenSearch/Elasticsearch (hot/warm/cold).
  • Graph / Vector Neo4j (knowledge graph, entity relationships). PGVector (semantic embeddings, similarity search).
  • AI Integration Langraph (agent orchestration). Claude API (primary LLM). Kimi K2. HITL pipeline gates.
  • Observability OpenTelemetry (trace propagation). Loki. Prometheus. Tempo. OPIK. Arize AI (ML pipeline observability).
  • Quality / CI SonarQube (quality gates). GitHub Actions (CI/CD). Playwright (E2E). Swagger/OpenAPI. ArgoCD.

Required Qualifications

  • 6+ years of backend engineering experience, with at least 2 years in a senior or lead capacity
  • Expert-level Python: FastAPI, Pydantic, asyncio, dataclasses — production code, not scripts
  • Production experience with Dagster or a comparable pipeline orchestrator (Airflow, Prefect) — Dagster strongly preferred
  • Deep PostgreSQL proficiency: query optimization, indexing, partitioning, HA configuration
  • Kafka or comparable streaming platform at meaningful throughput (100GB+/month minimum)
  • Redis operational experience: deduplication patterns, caching, pub/sub
  • Strong distributed systems fundamentals: idempotency, retry budgets, backpressure, exactly-once semantics
  • OpenTelemetry or comparable distributed tracing implementation — end-to-end trace ID propagation

Strong Advantage

  • Neo4j and Cypher — graph schema design, relationship modeling, query optimization
  • PGVector or comparable vector database for semantic similarity search
  • OpenSearch/Elasticsearch — hot/warm/cold tier management, index lifecycle policies, mapping design
  • Langraph, LangChain, or comparable agent orchestration framework
  • HITL pipeline design: human review gate implementation, feedback loop architecture, negative linking enforcement
  • Medallion data lake architecture at production scale
  • Data quality, lineage, and confidence scoring system design
  • Air-gapped deployment experience — all services must run fully offline

What We Offer

  • Senior-level remote compensation benchmarked globally
  • Architectural influence on a complex, multi-TB intelligence pipeline from day one
  • Exposure to graph databases, vector search, multi-agent AI orchestration, and OSINT data at scale
  • Mentorship opportunities and a clear path to Staff / Principal Engineer
  • A team that values depth, ownership, and directness