
How AI agents prepare clinical protocol review drafts with ThinkStack
AI agents can prepare a clinical protocol review draft by extracting content, retrieving reference documents, comparing selected design elements, and linking findings to their sources. ThinkTrends' CPAR (Clinical Protocol AI Review) proof of concept brings those steps together on ThinkStack. It targets the preparation work a clinical or regulatory review lead must coordinate before specialists assess the protocol. The supplied example shows configurations and draft outputs; experts remain responsible for interpreting guidance, checking findings, and approving any resulting action.
What slows protocol review preparation?
A protocol review lead coordinates questions about document completeness, statistical design, scientific rationale, and participant safety. Preparing that review means locating passages in the protocol, checking reference versions, and assembling evidence for several specialists. A finding needs enough context for the next reviewer to inspect it without repeating the search.
ICH M11 CeSHarP (Clinical Electronic Structured Harmonised Protocol) provides a harmonized structure for protocol information. A review checklist also needs to state which references apply and how to handle information supplied in other documents. CPAR uses a configured checklist to organize the draft.
Guidance, historical protocols, and biomedical terminology provide the prototype's reference material.
Statistical, scientific-rationale, and safety agents contribute distinct review perspectives.
Extract the content, compare the evidence, and assemble the review draft.
An expert determines whether each finding is supported, applicable, and worth acting on.
Which sources does CPAR use?
ThinkStack connects the model, review instructions, knowledge resources, and policies used by each agent. In the CPAR design, extraction, comparison, and report preparation share reference material. A review team must select and maintain that material for its intended use.
The prototype has three knowledge bases. ICH M11 & FDA Guidance contains selected review references; Historical Protocols contains comparison documents; Biomedical Ontologies contains terminology resources. Before a pilot, the review lead should approve the source set, record document versions, and define how outdated or conflicting references will be handled.

The prototype groups guidance, historical protocols, and terminology into separate knowledge bases that agents can share.

The CPAR system graphic shows extraction, comparison, and report preparation around shared knowledge and a human review handoff. The boundary depicts configured policy checks around the prototype.
ThinkStack's Knowledge Base Management documents source attribution, permission-aware retrieval, and document lifecycle controls. Policy & Guardrails applies configured rules, and Observability & Tracing records retrievals and tool calls. The Evaluation Harness supports rubrics and human calibration. These capabilities let the team examine both the draft and the execution behind it; deployment-specific access and retention still need confirmation.
When a finding is disputed, the team can inspect its retrieved passages and evaluate a revised agent against the same test cases. That gives the review lead a way to investigate whether the error came from the source set, retrieval, instructions, or evaluator.
How does CPAR prepare a review draft?
For a pilot, use a handoff like the one below to separate agent preparation from expert decisions. The existing CPAR example supplies draft artifacts; this table proposes the review responsibilities around them.
| Step | Agent action | Human responsibility | Output or control |
|---|---|---|---|
| Intake and extraction | Organize content and cite the source | Confirm versions and resolve missing or conflicting content | Checklist draft and exception list |
| Comparison | Retrieve comparators and propose findings | Check relevance, applicability, and severity | Source-linked comparison draft |
| Specialist analysis | Apply defined domain instructions | Accept, revise, or reject each finding | Attributable specialist findings |
| Report handoff | Assemble the Key Findings draft | Approve substantive conclusions and any downstream action | Reviewed packet with unresolved items identified |
A reviewer starts the interactive path by submitting a protocol and a review request. The extraction agent establishes the document identity and version, retrieves relevant material, and organizes content against the configured checklist. The draft should pair each Complete, Partial, or Missing status with the extracted evidence and its source location.
“Missing” means the agent did not locate the requested content in the reviewed material. The responsible specialist checks whether the content was overlooked, supplied elsewhere, or applicable to the trial. For a pilot, ambiguous or contradictory evidence should remain on an exception list until the reviewer resolves it.

The extraction agent combines a model, review instructions, an attached policy, and knowledge resources.
Ask the agent to review NCT04019964, organize the content, and cite the source evidence.
Use the selected guidance and terminology resources to support extraction and source linking.
Establish the title, NCT number, source document, and version before assessing the content.
Classify the 18 prototype categories as Complete, Partial, or Missing, with evidence for expert review.

The demonstration starts by establishing the protocol identity and source document.

The extraction screenshot shows the prototype's 18-category table. The table uses the prototype checklist labels.

The draft surfaces missing and partial content alongside strengths. Its section labels and severity judgments require expert verification.
The gap-analysis agent retrieves reference protocols and compares selected design elements. Its draft identifies what each protocol states and why the difference may deserve review. The reviewer checks whether the comparator is relevant before using that difference to support a finding.
The supplied example selected two melanoma checkpoint-inhibitor protocols, NCT02731729 and NCT03620019, for comparison with the prostate cancer protocol. The disease settings differ. A review lead should set comparator-selection criteria and check whether the trial populations, treatment context, and design support each comparison.

The draft names its comparators so a reviewer can assess their relevance before relying on the comparison.

The statistical comparison shows the agent's proposed findings and severity labels. A specialist must verify the evidence, trial context, and interpretation.

The comparison also reports strengths and recommendations. These are draft assessments to review alongside the comparator limitations.
A public historical protocol documents another study's design. Its presence in the knowledge base establishes neither regulatory acceptance nor a requirement for the study under review. The draft should identify whether a question comes from guidance or from a comparator difference.
The example raises questions about statistical justification, monitoring, and safety reporting. The assigned specialist checks the cited evidence, confirms applicability, and accepts, revises, or rejects each finding. Strengths and comparison limitations belong in the same review packet.
CPAR's coordinating agent has three specialist collaborators. The statistical agent examines sample-size justification and analysis methods; the scientific-rationale agent examines the rationale and endpoints; the safety agent examines reporting language and monitoring provisions. Human specialists validate the findings in their domains.

Shared resources let the coordinating agent bring distinct specialist instructions into the same review task.
ThinkStack supplies the delegation mechanism. The review team defines each specialist's remit, sources, and acceptance criteria. Findings need to remain attributable to the agent and evidence that produced them.
Adding a clinical-pharmacology or biomarker agent requires appropriate sources and representative test cases. The review lead should evaluate whether the added specialist finds useful issues without duplicating findings or increasing unnecessary review work.
ThinkStack's Policy & Guardrails supports content, topic, word, sensitive-information, and grounding rules. Teams select the rules, test their actions, and inspect enforcement records. CPAR's policy is intended to constrain review scope and handle sensitive content and unsupported output; the configured behavior needs testing against the pilot's input and failure cases.
Grounding checks compare output with retrieved material. A passage can pass that check and still be outdated or misapplied. The expert review therefore needs to assess the source and interpretation as well as the policy outcome.
The information requested could not be verified against
the knowledge base. Please ask a specific question about
the protocol content, M11 sections, or gap findings."
# A policy block does not establish regulatory accuracy.

Policy testing exposes configured rules and their behavior before publication.
Calibrate the model judge
CPAR's evaluation setup uses five criteria: section mapping, completeness, gap identification, citations, and regulatory language. ThinkStack documents custom rubrics, versioned expectations, regression runs, and blind human scoring to calibrate the model judge. A review lead can compare automated judgments with specialist assessments before relying on the scores. See the evaluation screen.

The displayed automated scores are 0.9, 0.9, and 1.0. The manual fields are empty, so this screen does not establish expert agreement or independently validated accuracy.

The evaluator-reasoning view makes automated judgments inspectable. Its explanations still need comparison with expert assessments.
Two QA agents also assess the draft outputs. Their judgments should be tested against known failure cases and expert decisions. Agent agreement alone supplies no independent reference answer.
The trace lets the team inspect the retrievals and tool calls behind a disputed finding. After changing the sources, instructions, or evaluator, the team should rerun the same cases and check for new errors as well as improvements.
The CPAR example includes three pipeline stages: extraction, gap analysis, and report generation. ThinkStack's code environment displays the source, dependencies, and deployment configuration. The prepared output is a Key Findings draft for reviewer assessment.
The three-stage pipeline graphic, retained as an illustration of the run sequence. Its status badges do not document a verified batch run.

The pipeline screen shows the implementation and deployment controls for the prototype. It is not evidence of production throughput or unattended batch reliability.
ThinkStack's Agentic ETL documents extraction, validation, and low-confidence review handling. Flow Orchestrator supports fixed node sequences, while Agent Workspaces supports dynamic delegation. CPAR's illustrated pipeline follows a fixed preparation sequence; a team should choose its orchestration route based on the actual handoffs and exception paths.
Scheduled preparation is a possible deployment path. Before enabling it, test the trigger, failed-stage handling, source access, and report destination. Define who owns incomplete runs and prevent an unfinished packet from entering the normal review handoff.
What does the supplied demonstration show?
The supplied CPAR material shows a retrospective review of NCT04019964, a prostate cancer protocol, with selected melanoma comparators. Its screenshots document configurations and AI-generated drafts. The public protocol itself has not been independently re-reviewed for this article.
The supplied demonstration reports 26 potential findings across seven configured domains: statistical/biometrics, regulatory/administrative, clinical/medical, pharmacovigilance, biomarker/companion diagnostics, clinical pharmacology, and chemistry, manufacturing, and controls (CMC). The agent assigned seven Critical, 14 Major, and five Minor labels. See the supplied draft output.
Agent-assigned severity
Draft scores by review domain
- Statistical / BiometricsAgent flags statistical justification, monitoring, and stopping-rule questions
- Regulatory / AdminAgent flags two checklist categories and a retention question
- Clinical / MedicalAgent proposes clinical review questions
- PharmacovigilanceAgent labels two findings Critical
- Biomarker / CDxAgent flags specification questions
- Clinical PharmacologyAgent reports fewer proposed gaps
- CMCAgent labels one drug-destruction finding Minor
The charts describe the draft's classifications. Specialist assessment is needed to establish which findings are supported and how their severity should be interpreted. The supplied material includes no confirmed error rate, reviewer-time baseline, or measured reduction in review effort.
The extraction draft flags financing/insurance and publication-policy content. ICH E6(R3), Appendix B addresses those topics. A specialist must check the applicable references and full document set before treating either absence as a deficiency.
The prototype's 18-category checklist uses M11-prefixed labels that differ from the final ICH M11 template. Reconcile the checklist and reference versions before evaluating it for an intended use.
Prototype checklist output
Evaluate a pilot using representative protocols and expert-reviewed findings. Measure missed issues, unsupported findings, citation correctness, severity agreement, active preparation time, expert review time, rework, and operating cost. Include poor-quality documents and weak comparators so the evaluation tests the exception path as well as ordinary cases.
What does a pilot require?
A pilot needs readable protocols, an approved source set, authorized access, and a named owner for each exception. Confirm that reviewers can open the cited sources and identify the version used. Define the handoff format and how reviewers record corrections before connecting the workflow to another system.
Scanned protocols require extraction testing on representative pages, tables, and layouts. ThinkStack documents OCR capabilities, but the CPAR example uses text-based PDFs. Review the extracted evidence before applying the comparison workflow to scans.
Access restrictions and retrieval failures need explicit handling. If the agent cannot retrieve a required source, the draft should identify the limitation for reviewer resolution. Test sensitive-information policies, retention, and the report delivery path in the intended environment.
Define the review task before starting a pilot
Select one review task, an approved set of protocols, and the specialists who will assess the outputs. Record today's preparation and review effort, then agree on acceptance criteria for missing issues, unsupported findings, and source accuracy.
Use those criteria to compare the CPAR draft with the existing process. Proceed to a broader pilot only when the review lead can account for the findings, exception handling, and total effort required to produce an accepted review packet.