LLM tooling for trials
Eligibility screening, FDA audit documentation workflow, and speech-battery stimulus generation.
What
Three production tools built with large language models, inside a clinical-stage biotech:
- Eligibility screening — automated assessment of candidate patients against a protocol’s inclusion and exclusion criteria.
- Audit-level documentation — a workflow that produces trial documentation to the standard an FDA audit requires.
- A speech task battery — a spoken assessment whose stimuli are GenAI-generated, now running in a live clinical trial.
Why it mattered
Clinical trials are slow in specific, unglamorous places: reading a chart against twenty criteria, writing the paper trail that makes a result defensible, and building task material that is novel enough not to be memorised between visits. None of that is the science, and all of it sets the pace. It is also the exact shape of work language models handle well — provided someone who knows the protocol defines what “correct” means and checks that it holds.
My role
Built and shipped all three, working across clinical, engineering, and drug-development teams on the requirements and on evaluating output against them.
Outcome
All three reached production use. The speech task battery, with its generated stimuli, is running in a live clinical trial.