The problem
Regulatory submissions are assembled by hand from clinical study reports by expert medical writers. The work is slow, expensive and hard to scale — and every sentence has to be defensible against the underlying data, which rules out unconstrained generation.
Engagement detail
- CLIENT
- Global pharmaceutical enterprise
- INDUSTRY
- Pharmaceuticals & life sciences
- DISCIPLINE
- Generative AI
PythonLangChainGPT (OpenAI)RAGPrompt engineeringAWS EC2
What we built
- Designed a multi-step LLM pipeline in LangChain that generates each section of the regulatory document from raw clinical trial data.
- Used a Retrieval-Augmented Generation architecture to pull context from across multiple clinical study reports, keeping generated content factually anchored and internally coherent.
- Wrote specialised prompts per section — study design, results, safety profile — so each one meets its own regulatory conventions and scientific register.
- Orchestrated sequenced LLM calls for structured, long-form content generation that mirrors how an expert medical writer builds a document.
- Ran the output through domain-expert validation loops to verify accuracy before it entered the submission workflow.
Available for new engagements
Have a problem shaped like this one?
We will tell you what transfers from this engagement to yours, what does not, and what we would do differently now.