Founding Engineer, ML Infrastructure & Evaluation
A gate is only trusted if it is measured. You build the evaluation harness, the tuning pipelines, and the deployment rails for every model behind it.
The problem you own
Cull self-hosts and tunes its speech and judgment models, and the roadmap moves toward governance-tuned models of our own: diarization and gating optimized for capture decisions rather than transcription accuracy. Someone owns the machinery: training and tuning pipelines, the evaluation harness that makes the gate's accuracy claims credible to buyers and courts, deployment, observability, and a GPU bill that stays honest.
What you will do
Build fine-tuning and continual-evaluation pipelines for ASR, diarization, and gate-classification models.
Own the evaluation harness as a product: per-speaker accuracy, disposition correctness, latency distributions.
Own deployment and observability across the fleet, from cloud today toward on-device and in-tenant builds.
Make cost per governed minute a first-class metric the founders can quote.
What great looks like
You have run ML infrastructure for a small research team and kept it fast, measured, and affordable, and researchers voluntarily said thank you.
Stack and signals
Python, PyTorch, orchestration you can defend, GPUs and their invoices, metrics as a love language.
Compensation
Cash is deliberately founder-grade; the equity is the point, and it is real. You will work directly with the founding technical leadership, in the codebase, from week one.
How we hire
One conversation about your best work, one deep-dive on a system or a deal you owned, one paid working session on a real Cull problem, one founder conversation. Usually under two weeks.
jobs@cull.ai · Subject: Founding Engineer, ML Infrastructure & Evaluation. Send evidence, not adjectives.