About the Role
We are building ID1X — a large-scale identity intelligence platform that ingests, deduplicates, normalizes, and fuses data from dozens of OSINT vendors into a unified identity graph. The data pipeline processes terabytes of structured and unstructured data monthly through a Medallion architecture (Bronze → Silver → Gold), feeding downstream AI agents, a knowledge graph (Neo4j), a vector database (PGVector), and a unified search engine (OpenSearch/M-Mind).
This is a senior individual contributor role. You will architect and build the backend services and data pipelines that power the entire platform — and you will mentor junior engineers on engineering standards, data modeling, and system design. The role has meaningful architectural influence from day one.
Key Responsibilities
- Design and implement Dagster-orchestrated data pipelines across the full Medallion stack: Bronze (raw ingestion) → Silver (normalization, dedup, enrichment) → Gold (identity graph population, vector embedding generation)
- Build and own the vendor data middleware layer: normalize OSINT vendor outputs (District4, IntelliX, Bright Data, Social Links, Constella) to a canonical internal schema with mandatory fields, source quality scores, and confidence weighting
- Implement deduplication and identity resolution logic: Redis-backed fingerprinting for idempotent ingestion, Confluent Kafka for high-throughput streaming at 1TB+/month inbound
- Design and optimize PostgreSQL schemas for job state management, identity tracking, and audit logging; implement read replicas, partitioning, and migration strategies
- Build and maintain the ACP (API Control Plane): query planner, normalization layer, adapter pattern for provider-agnostic data access, Swagger documentation, P99 ≤ 20ms performance baseline
- Integrate with Neo4j knowledge graph: design entity/relationship schemas, Cypher query optimization, graph enrichment pipeline
- Integrate with PGVector for semantic embedding storage and similarity search; coordinate with AI/ML engineers on embedding generation models
- Implement OpenTelemetry trace propagation (Trace ID / Flow ID / UUID) end-to-end across all backend services — critical for support and incident debugging
- Enforce backend engineering standards: code review, SonarQube quality gates, test coverage, CI/CD discipline via GitHub Actions
- Mentor junior and mid-level engineers; contribute to architecture decision records (ADRs) and system design documentation
Tech Stack You'll Own
- Language Python (primary) — expert level. FastAPI. Pydantic v2. asyncio. Type annotations throughout.
- Pipeline Dagster (orchestration). Medallion architecture (Bronze/Silver/Gold). Provider-agnostic ingestion design.
- Messaging Confluent Kafka (streaming, 1TB+/month). RabbitMQ (job queuing). Redis (dedup fingerprinting, caching).
- Storage PostgreSQL (HA, partitioning, read replicas). MinIO (object storage).
- OpenSearch/Elasticsearch (hot/warm/cold).
- Graph / Vector Neo4j (knowledge graph, entity relationships). PGVector (semantic embeddings, similarity search).
- AI Integration Langraph (agent orchestration). Claude API (primary LLM). Kimi K2. HITL pipeline gates.
- Observability OpenTelemetry (trace propagation). Loki. Prometheus. Tempo. OPIK. Arize AI (ML pipeline observability).
- Quality / CI SonarQube (quality gates). GitHub Actions (CI/CD). Playwright (E2E). Swagger/OpenAPI. ArgoCD.
Required Qualifications
- 6+ years of backend engineering experience, with at least 2 years in a senior or lead capacity
- Expert-level Python: FastAPI, Pydantic, asyncio, dataclasses — production code, not scripts
- Production experience with Dagster or a comparable pipeline orchestrator (Airflow, Prefect) — Dagster strongly preferred
- Deep PostgreSQL proficiency: query optimization, indexing, partitioning, HA configuration
- Kafka or comparable streaming platform at meaningful throughput (100GB+/month minimum)
- Redis operational experience: deduplication patterns, caching, pub/sub
- Strong distributed systems fundamentals: idempotency, retry budgets, backpressure, exactly-once semantics
- OpenTelemetry or comparable distributed tracing implementation — end-to-end trace ID propagation
Strong Advantage
- Neo4j and Cypher — graph schema design, relationship modeling, query optimization
- PGVector or comparable vector database for semantic similarity search
- OpenSearch/Elasticsearch — hot/warm/cold tier management, index lifecycle policies, mapping design
- Langraph, LangChain, or comparable agent orchestration framework
- HITL pipeline design: human review gate implementation, feedback loop architecture, negative linking enforcement
- Medallion data lake architecture at production scale
- Data quality, lineage, and confidence scoring system design
- Air-gapped deployment experience — all services must run fully offline
What We Offer
- Senior-level remote compensation benchmarked globally
- Architectural influence on a complex, multi-TB intelligence pipeline from day one
- Exposure to graph databases, vector search, multi-agent AI orchestration, and OSINT data at scale
- Mentorship opportunities and a clear path to Staff / Principal Engineer
- A team that values depth, ownership, and directness