How AI agents prepare clinical protocol review drafts with ThinkStack

How AI agents prepare clinical protocol review drafts with ThinkStack

CPAR shows how ThinkStack organizes protocol evidence, specialist analysis, and review drafts. See the human checks and requirements for a pilot.
Tirna Deb, Max Holmes
October 9, 2026
Share on TwitterShare on LinkedIn

AI agents can prepare a clinical protocol review draft by extracting content, retrieving reference documents, comparing selected design elements, and linking findings to their sources. ThinkTrends' CPAR (Clinical Protocol AI Review) proof of concept brings those steps together on ThinkStack. It targets the preparation work a clinical or regulatory review lead must coordinate before specialists assess the protocol. The supplied example shows configurations and draft outputs; experts remain responsible for interpreting guidance, checking findings, and approving any resulting action.


What slows protocol review preparation?

A protocol review lead coordinates questions about document completeness, statistical design, scientific rationale, and participant safety. Preparing that review means locating passages in the protocol, checking reference versions, and assembling evidence for several specialists. A finding needs enough context for the next reviewer to inspect it without repeating the search.

ICH M11 CeSHarP (Clinical Electronic Structured Harmonised Protocol) provides a harmonized structure for protocol information. A review checklist also needs to state which references apply and how to handle information supplied in other documents. CPAR uses a configured checklist to organize the draft.

Multiple Knowledge Bases

Guidance, historical protocols, and biomedical terminology provide the prototype's reference material.

Specialist Agent Collaborators

Statistical, scientific-rationale, and safety agents contribute distinct review perspectives.

Consistent Pipeline Stages

Extract the content, compare the evidence, and assemble the review draft.

Accountable Review Handoff

An expert determines whether each finding is supported, applicable, and worth acting on.


Which sources does CPAR use?

ThinkStack connects the model, review instructions, knowledge resources, and policies used by each agent. In the CPAR design, extraction, comparison, and report preparation share reference material. A review team must select and maintain that material for its intended use.

ENTRY POINTSChatReviewer asks, ad hocCollaboratePipelineConfigurable executionAgentic ETLinvokesinvokesONE AGENT LAYERProtocol Extractionprimary · 18 review categoriesGap Analysis7 review domains vs. comparators3 Collaborator Agentsstatistical · rationale · safety2 QA Agentsgrade the other agents' outputConfigured policy: runtime enforcementgrounding checks against retrieved contentGROUNDED INICH M11 & FDACeSHarP template · E6(R3)Historical Protocolsindexed by area & drug classBiomedical OntologiesBRIDG · FDA data standardsKey Findings: DRAFTExpert reviewer validates and correctsGOVERNANCEEvaluation Studio5 criteria · human calibrationTrace Explorerevery call, every retrievalPromptsversioned instruction setsscores · traces · versionsflow back into theagent layerALSO ON THE PLATFORMMCP Servers: live database lookupsConnected Apps: Slack, Jira, WorkspaceSecrets Manager: stores credentialsFlow Orchestrator: node-based flowsSearch: across connected sourcesconfiguration and access apply
CPAR architecture: shared agent and knowledge resources support interactive preparation and pipeline execution. Governance supplies policy checks, evaluation, and trace inspection. The additional platform capabilities depend on configured tools and access.

The prototype has three knowledge bases. ICH M11 & FDA Guidance contains selected review references; Historical Protocols contains comparison documents; Biomedical Ontologies contains terminology resources. Before a pilot, the review lead should approve the source set, record document versions, and define how outdated or conflicting references will be handled.

ThinkStack Knowledge Bases screen showing the three CPAR knowledge bases, Biomedical Ontologies, Historical Protocols, and ICH M11 and FDA Guidance, all with Available status and recent sync timestamps

The prototype groups guidance, historical protocols, and terminology into separate knowledge bases that agents can share.

CPAR system architecture diagram: three stages, Protocol Extraction Agent, Gap Analysis Agent, Automated Report Pipeline, inside a compliance guardrail boundary, over three knowledge bases, with Evaluation Studio scoring and a human-in-the-loop CDER reviewer

The CPAR system graphic shows extraction, comparison, and report preparation around shared knowledge and a human review handoff. The boundary depicts configured policy checks around the prototype.

ThinkStack's Knowledge Base Management documents source attribution, permission-aware retrieval, and document lifecycle controls. Policy & Guardrails applies configured rules, and Observability & Tracing records retrievals and tool calls. The Evaluation Harness supports rubrics and human calibration. These capabilities let the team examine both the draft and the execution behind it; deployment-specific access and retention still need confirmation.

When a finding is disputed, the team can inspect its retrieved passages and evaluate a revised agent against the same test cases. That gives the review lead a way to investigate whether the error came from the source set, retrieval, instructions, or evaluator.


How does CPAR prepare a review draft?

For a pilot, use a handoff like the one below to separate agent preparation from expert decisions. The existing CPAR example supplies draft artifacts; this table proposes the review responsibilities around them.

Proposed CPAR pilot responsibilities
StepAgent actionHuman responsibilityOutput or control
Intake and extractionOrganize content and cite the sourceConfirm versions and resolve missing or conflicting contentChecklist draft and exception list
ComparisonRetrieve comparators and propose findingsCheck relevance, applicability, and severitySource-linked comparison draft
Specialist analysisApply defined domain instructionsAccept, revise, or reject each findingAttributable specialist findings
Report handoffAssemble the Key Findings draftApprove substantive conclusions and any downstream actionReviewed packet with unresolved items identified

A reviewer starts the interactive path by submitting a protocol and a review request. The extraction agent establishes the document identity and version, retrieves relevant material, and organizes content against the configured checklist. The draft should pair each Complete, Partial, or Missing status with the extracted evidence and its source location.

“Missing” means the agent did not locate the requested content in the reviewed material. The responsible specialist checks whether the content was overlooked, supplied elsewhere, or applicable to the trial. For a pilot, ambiguous or contradictory evidence should remain on an exception list until the reviewer resolves it.

ThinkStack agent configuration page for CPAR-Agent-Protocol Extraction Agent showing name, description, foundation model Claude Sonnet 4.6, bound guardrail, system prompt enforcing the ICH M11 taxonomy, and three attached knowledge bases

The extraction agent combines a model, review instructions, an attached policy, and knowledge resources.

1
Submit the protocol

Ask the agent to review NCT04019964, organize the content, and cite the source evidence.

2
Retrieve the reference material

Use the selected guidance and terminology resources to support extraction and source linking.

3
Confirm the document identity

Establish the title, NCT number, source document, and version before assessing the content.

4
Return the configured checklist

Classify the 18 prototype categories as Complete, Partial, or Missing, with evidence for expert review.

ThinkStack chat showing the reviewer's question about NCT04019964 and the agent's structured Protocol Identity table with full title, protocol ID, PI, sponsor, source document and ICH M11 standard

The demonstration starts by establishing the protocol identity and source document.

ThinkStack chat output showing the ICH M11 section analysis table for sections M11-6 through M11-18 with Complete, Partial and Missing flags, and a summary table reading 13 Complete 72 percent, 3 Partial 17 percent, 2 Missing 11 percent

The extraction screenshot shows the prototype's 18-category table. The table uses the prototype checklist labels.

ThinkStack chat output titled Key Findings and Recommendations, organised into Critical Gaps showing missing Financing and Insurance and Publication Policy sections, Gaps to Address, and Protocol Strengths

The draft surfaces missing and partial content alongside strengths. Its section labels and severity judgments require expert verification.

The gap-analysis agent retrieves reference protocols and compares selected design elements. Its draft identifies what each protocol states and why the difference may deserve review. The reviewer checks whether the comparator is relevant before using that difference to support a finding.

The supplied example selected two melanoma checkpoint-inhibitor protocols, NCT02731729 and NCT03620019, for comparison with the prostate cancer protocol. The disease settings differ. A review lead should set comparator-selection criteria and check whether the trial populations, treatment context, and design support each comparison.

ThinkStack chat showing the FDA Review Domain Gap Analysis header, naming both retrieved melanoma reference protocols and beginning Domain 1 Statistical and Biometrics

The draft names its comparators so a reviewer can assess their relevance before relying on the comparison.

ThinkStack chat output table for Domain 1 Statistical and Biometrics comparing six elements across the prostate protocol and the melanoma reference standard with Critical and Major severity ratings

The statistical comparison shows the agent's proposed findings and severity labels. A specialist must verify the evidence, trial context, and interpretation.

ThinkStack chat output listing four areas where the prostate protocol exceeds the melanoma reference standard, followed by a bottom-line recommendation naming the two most urgent fixes

The comparison also reports strengths and recommendations. These are draft assessments to review alongside the comparator limitations.

A public historical protocol documents another study's design. Its presence in the knowledge base establishes neither regulatory acceptance nor a requirement for the study under review. The draft should identify whether a question comes from guidance or from a comparator difference.

The example raises questions about statistical justification, monitoring, and safety reporting. The assigned specialist checks the cited evidence, confirms applicability, and accepts, revises, or rejects each finding. Strengths and comparison limitations belong in the same review packet.

CPAR's coordinating agent has three specialist collaborators. The statistical agent examines sample-size justification and analysis methods; the scientific-rationale agent examines the rationale and endpoints; the safety agent examines reporting language and monitoring provisions. Human specialists validate the findings in their domains.

ThinkStack agent resources panel showing three knowledge bases and three collaborator agents, Statistical Reviewer, Scientific Rationale Reviewer and Safety Reviewer, each tagged with an orange Collaborator badge, with an Add Collaborators button

Shared resources let the coordinating agent bring distinct specialist instructions into the same review task.

ThinkStack supplies the delegation mechanism. The review team defines each specialist's remit, sources, and acceptance criteria. Findings need to remain attributable to the agent and evidence that produced them.

Adding a clinical-pharmacology or biomarker agent requires appropriate sources and representative test cases. The review lead should evaluate whether the added specialist finds useful issues without duplicating findings or increasing unnecessary review work.

ThinkStack's Policy & Guardrails supports content, topic, word, sensitive-information, and grounding rules. Teams select the rules, test their actions, and inspect enforcement records. CPAR's policy is intended to constrain review scope and handle sensitive content and unsupported output; the configured behavior needs testing against the pilot's input and failure cases.

Grounding checks compare output with retrieved material. A passage can pass that check and still be outdated or misapplied. The expert review therefore needs to assess the source and interpretation as well as the policy outcome.

Content Filters· Sensitive Information· Topic Filters· Word Filters· Contextual Grounding
Example policy-block message from the prototype "This response was blocked to ensure regulatory accuracy.
The information requested could not be verified against
the knowledge base. Please ask a specific question about
the protocol content, M11 sections, or gap findings."

# A policy block does not establish regulatory accuracy.
ThinkStack compliance guardrail configuration page for CPAR-Guardrail-Clinical Protocol Compliance showing the blocked output message and the five collapsible policy sections: Content Filters, Sensitive Information, Topic Filters, Word Filters and Contextual Grounding, with a Test Guardrail panel

Policy testing exposes configured rules and their behavior before publication.

Calibrate the model judge

CPAR's evaluation setup uses five criteria: section mapping, completeness, gap identification, citations, and regulatory language. ThinkStack documents custom rubrics, versioned expectations, regression runs, and blind human scoring to calibrate the model judge. A review lead can compare automated judgments with specialist assessments before relying on the scores. See the evaluation screen.

ThinkStack Evaluation Studio scores screen showing 3 total cases, 3 high scores, 0 low scores, a blind manual review banner, and agent scores of 0.9, 0.9 and 1.0 with empty manual score fields

The displayed automated scores are 0.9, 0.9, and 1.0. The manual fields are empty, so this screen does not establish expert agreement or independently validated accuracy.

ThinkStack Evaluation Studio detail panel showing criterion-by-criterion reasoning for the 1.0 case, with Section Mapping Accuracy and Content Completeness verification checklists

The evaluator-reasoning view makes automated judgments inspectable. Its explanations still need comparison with expert assessments.

Two QA agents also assess the draft outputs. Their judgments should be tested against known failure cases and expert decisions. Agent agreement alone supplies no independent reference answer.

The trace lets the team inspect the retrievals and tool calls behind a disputed finding. After changing the sources, instructions, or evaluator, the team should rerun the same cases and check for new errors as well as improvements.

The CPAR example includes three pipeline stages: extraction, gap analysis, and report generation. ThinkStack's code environment displays the source, dependencies, and deployment configuration. The prepared output is a Key Findings draft for reviewer assessment.

DagCPAR-Pipeline-Protocol-Review
Dag RunIllustrative run layout
stage_1_protocol_extraction Maps the 18 prototype review categories · flags status · cites evidence ✓ illustrated success
stage_2_gap_analysis Retrieves comparators · compares 7 review domains · proposes severity ✓ illustrated success
generate_review_report Combines both stages into one structured Key Findings DRAFT for the reviewer ✓ illustrated success
Prototype pipelineAirflow · containerised tasksTriggers require deployment testingConfigured policy checksOutput: review DRAFT

The three-stage pipeline graphic, retained as an illustration of the run sequence. Its status badges do not document a verified batch run.

ThinkStack pipeline build screen showing the CPAR Clinical Protocol AI Review Pipeline with Ready and Deployed badges, a file tree containing main.py, pipeline_logic.py, Dockerfile, README.md and requirements.txt, the generated pipeline code, and an AI Assistant panel explaining recent changes to environment variables

The pipeline screen shows the implementation and deployment controls for the prototype. It is not evidence of production throughput or unattended batch reliability.

ThinkStack's Agentic ETL documents extraction, validation, and low-confidence review handling. Flow Orchestrator supports fixed node sequences, while Agent Workspaces supports dynamic delegation. CPAR's illustrated pipeline follows a fixed preparation sequence; a team should choose its orchestration route based on the actual handoffs and exception paths.

Scheduled preparation is a possible deployment path. Before enabling it, test the trigger, failed-stage handling, source access, and report destination. Define who owns incomplete runs and prevent an unfinished packet from entering the normal review handoff.


What does the supplied demonstration show?

The supplied CPAR material shows a retrospective review of NCT04019964, a prostate cancer protocol, with selected melanoma comparators. Its screenshots document configurations and AI-generated drafts. The public protocol itself has not been independently re-reviewed for this article.

The supplied demonstration reports 26 potential findings across seven configured domains: statistical/biometrics, regulatory/administrative, clinical/medical, pharmacovigilance, biomarker/companion diagnostics, clinical pharmacology, and chemistry, manufacturing, and controls (CMC). The agent assigned seven Critical, 14 Major, and five Minor labels. See the supplied draft output.

Agent-assigned severity

7
14
5
7 Critical · 26.9% · agent-assigned highest severity
14 Major · 53.8% · agent-assigned major severity
5 Minor · 19.2% · agent-assigned minor severity
Draft severity distribution: 26 potential findings, categorized by the agent. The chart describes draft labels; expert validation is needed to confirm each finding and severity.

Draft scores by review domain

Statistical / Biometrics
1.0
Regulatory / Admin
3.8
Clinical / Medical
5.0
Pharmacovigilance
5.8
Biomarker / CDx
7.0
Clinical Pharmacology
8.5
CMC
9.8
0510
  • Statistical / BiometricsAgent flags statistical justification, monitoring, and stopping-rule questions
  • Regulatory / AdminAgent flags two checklist categories and a retention question
  • Clinical / MedicalAgent proposes clinical review questions
  • PharmacovigilanceAgent labels two findings Critical
  • Biomarker / CDxAgent flags specification questions
  • Clinical PharmacologyAgent reports fewer proposed gaps
  • CMCAgent labels one drug-destruction finding Minor
Draft domain scores on the prototype's 0–10 scale, with higher values indicating fewer proposed gaps. The displayed scores use the configured severity penalties; their clinical interpretation and underlying findings require specialist review.

The charts describe the draft's classifications. Specialist assessment is needed to establish which findings are supported and how their severity should be interpreted. The supplied material includes no confirmed error rate, reviewer-time baseline, or measured reduction in review effort.

The extraction draft flags financing/insurance and publication-policy content. ICH E6(R3), Appendix B addresses those topics. A specialist must check the applicable references and full document set before treating either absence as a deficiency.

The prototype's 18-category checklist uses M11-prefixed labels that differ from the final ICH M11 template. Reconcile the checklist and reference versions before evaluating it for an intended use.

Prototype checklist output

Complete: 13 (72%) Partial: 3 (17%) Missing: 2 (11%)
M11-1Title Page● COMPLETE
M11-2Protocol Amendment History● COMPLETE
M11-3Protocol Synopsis● COMPLETE
M11-4Introduction / Background● COMPLETE
M11-5Study Objectives & Endpoints● COMPLETE
M11-6Study Design● COMPLETE
M11-7Study Population● COMPLETE
M11-8Investigational Treatments● COMPLETE
M11-9Schedule of Activities● COMPLETE
M11-10Efficacy Assessments● COMPLETE
M11-11Safety Assessments● COMPLETE
M11-12Adverse Events / Reporting◐ PARTIAL
M11-13Statistical Considerations◐ PARTIAL
M11-14Ethical & Regulatory● COMPLETE
M11-15Data Handling & Records● COMPLETE
M11-16Financing and Insurance✕ MISSING
M11-17Publication Policy✕ MISSING
M11-18Appendices◐ PARTIAL
Prototype classifications for NCT04019964: 13 Complete, three Partial, and two Missing. The M11-prefixed labels reproduce the original checklist, rather than the final M11 template numbering. These are AI-generated assessments.

Evaluate a pilot using representative protocols and expert-reviewed findings. Measure missed issues, unsupported findings, citation correctness, severity agreement, active preparation time, expert review time, rework, and operating cost. Include poor-quality documents and weak comparators so the evaluation tests the exception path as well as ordinary cases.

A finding should identify the source passage, the question it raises, and the specialist responsible for resolving it.

What does a pilot require?

A pilot needs readable protocols, an approved source set, authorized access, and a named owner for each exception. Confirm that reviewers can open the cited sources and identify the version used. Define the handoff format and how reviewers record corrections before connecting the workflow to another system.

Scanned protocols require extraction testing on representative pages, tables, and layouts. ThinkStack documents OCR capabilities, but the CPAR example uses text-based PDFs. Review the extracted evidence before applying the comparison workflow to scans.

Access restrictions and retrieval failures need explicit handling. If the agent cannot retrieve a required source, the draft should identify the limitation for reviewer resolution. Test sensitive-information policies, retention, and the report delivery path in the intended environment.

Define the review task before starting a pilot

Select one review task, an approved set of protocols, and the specialists who will assess the outputs. Record today's preparation and review effort, then agree on acceptance criteria for missing issues, unsupported findings, and source accuracy.

Use those criteria to compare the CPAR draft with the existing process. Proceed to a broader pilot only when the review lead can account for the findings, exception handling, and total effort required to produce an accepted review packet.