Product: Identify
Identify: LLM-based EHR screening with the evidence attached
Identify is the first stage of Bond Health's recruitment workflow. Bond reads a site's EHR records, from structured fields to clinical notes, imaging data and other unstructured documents, against a study's inclusion and exclusion criteria, ranks the patients most likely to qualify, and shows the chart evidence behind each criterion decision.[1],[2] Coordinators review a ranked, explained list instead of opening charts one at a time.
What does Bond read in the EHR?
Bond connects to all the major EHRs, including Epic, Oracle Health (Cerner), MEDITECH, athenahealth, eClinicalWorks, NextGen, Veradigm and OncoEMR,[2] through FHIR R4 APIs, HL7 v2 feeds where applicable, or an integration partner.[1] The integrations page lists what each EHR needs from site IT. What Bond reads depends on what the site's connection exposes:
- Structured data: diagnoses and problem lists, medications and prescriptions, lab results, vitals, procedures and demographics. These settle criteria like age, a lab threshold or a current prescription.
- Clinical notes and reports: progress notes, history and physicals, discharge summaries, and pathology and imaging narratives. All five are among the eight clinical note types in version 1 of the federal USCDI data standard.[11]
- Imaging data and other documents: Identify also uses imaging data and other unstructured documents, including pathology, radiology and molecular reports.[2]
Notes matter because many criteria are never coded; the answer sits in unstructured clinical data. For a heart failure trial at Mass General Brigham, the team behind the RECTIFIER screening tool found that structured EHR data could determine 5 of 6 inclusion criteria but only 5 of 17 exclusion criteria. The rest required manual chart review.[4]
How does Bond decide whether a patient meets a criterion?
Bond's matching pipeline, described in its August 2026 technical report (an internal preprint), runs in three stages. Each language-model stage receives an evidence packet: the clinical concepts involved, their synonyms, their mappings across coding systems and their hierarchy relations, each tagged with where it came from.[3]
- 1Relevance. The first pass asks whether the study concerns this patient's condition at all.
- 2Criterion-level evaluation. Each inclusion and exclusion criterion is checked against the patient's record and the graph-grounded evidence, and the evidence used is kept with the decision.
- 3Ranking. Criterion results are combined into graded ranking signals. Scoring and gating follow deterministic rules rather than free-form model output.[3]
The terminology layer lets one check catch the many ways a chart says the same thing. Bond's ClinText Graph holds 3,270,078 nodes drawn from 18 biomedical terminologies, including SNOMED CT, RxNorm, LOINC, ICD-10-CM and MeSH. Tested on 51,055 eligibility criteria from ClinicalTrials.gov, it resolved at least one exact multi-token concept in 82.0 percent of criteria, against 5.7 percent for ICD-10-CM alone.[3]
How are the criteria configured and validated?
Screening starts from the study's own inclusion and exclusion criteria. Setting them up is part of the workflow configuration covered by Bond's platform fee, and it happens inside the implementation plan, which takes 48 hours for full EHR integration depending on the EHR, IT review and interface method.[1]
- 1
Settle ambiguous criteria
Phrases like "adequate organ function" need the thresholds the sponsor actually uses. The PI and study team decide those interpretations, not the software, and write them down.
- 2
Resolve criteria to concepts
Bond's terminology graph maps each criterion to concepts, synonyms and codes across vocabularies.[3] State time windows explicitly so that an old lab or a resolved diagnosis does not count.
- 3
Check a sample against the chart
Before relying on the ranked list, coordinators check a sample of criterion decisions against the evidence each one cites. Spend the most time on criteria that depend on dates or on combining several findings.
- 4
Keep measuring after launch
Matching accuracy, screen-failure signals and coordinator hours saved are among the outcomes Bond reports, and decisions stay in the audit trail.[1] When an amendment changes eligibility, repeat the sample check on the changed criteria.
What does the coordinator see when reviewing matches?
Coordinators work from a ranked list in Bond's real-time dashboard. Each candidate comes with the criteria, the decision on each, and the note, lab or medication behind it, so the coordinator checks that evidence rather than re-reading the whole record.[1]
- Gaps the chart cannot fill. Some criteria, such as willingness to use contraception, are rarely documented. Those still have to be asked in pre-screening, by the voice and SMS/text agents in Engage or by staff.
- Audit trail. Decisions and their rationale are recorded, which helps when a monitor asks why a patient was or was not approached.
How the output fits the coordinator's workflow matters as much as the model. In a 2026 randomized evaluation of 355 oncology charts from a community practice, AI assistance raised coordinators' chart-level accuracy from 71.1 to 76.5 percent but did not shorten full chart abstraction: 37.4 versus 37.8 minutes per chart.[9] A 2014 systematic review of 79 recruitment support systems concluded that success depends more on workflow integration than on sophisticated algorithms.[10]
Reduced our chart review time significantly while improving the quality of patients we bring in for screening.
How accurate is LLM-based eligibility screening?
Bond publishes two kinds of numbers: a matching accuracy figure on its website, and benchmark results from a technical report that is an internal preprint, not a peer-reviewed paper.[1],[3] Independent results are listed alongside.
90%+[1]
matching accuracy for eligibility screening, as reported by Bond
0.9312[3]
micro F1 on the held-out n2c2 2018 cohort selection set
82.0%[3]
of 51,055 trial criteria resolved to an exact concept, vs 5.7% with ICD-10-CM alone
| System and study | Setting | Reported result |
|---|---|---|
| Bond (internal preprint, 2026) | Held-out n2c2 2018 cohort selection set, organizers' scorer | 0.9312 overall micro F1[3] |
| Bond (internal preprint, 2026) | TREC Clinical Trials 2021 and 2022 fixed judged subsets | nDCG@10 up 0.1315 and 0.1316 over TrialGPT's criterion-count ranking[3] |
| n2c2 2018 shared task (JAMIA, 2019) | 288 patient records, 13 criteria, 47 teams | Best system micro F1 of 0.91, rule-based[7] |
| TrialGPT, NIH (Nature Communications, 2024) | 1,015 patient-criterion pairs, manual evaluation | 87.3% criterion-level accuracy; screening time down 42.6% in a user study[6] |
| RECTIFIER, Mass General Brigham (2024) | Test set of 1,894 heart failure patients, scored against a blinded expert clinician | Sensitivity 92.3% and specificity 93.9%, vs 90.1% and 83.6% for trained study staff[4] |
| RECTIFIER randomized trial, Mass General Brigham (JAMA, 2025) | 4,476 patients randomized to AI-assisted or manual prescreening | 20.4% vs 12.7% found eligible; 35 vs 19 enrollments[5] |
Benchmark sets use de-identified or synthetic records. They show what a method can do under test conditions, not what it will do on your protocol and EHR.
How does manual chart review compare with Bond screening?
In a 2012 prospective study at Virginia Commonwealth University's cancer center, the largest share of eligibility evaluations (35.8 percent) took 10 to 30 minutes, and more than 10 percent took 2 to 4 hours. Finding, screening and enrolling one patient took an average of 3.4 to 8.8 staff hours, depending on study phase.[8]
| Task | Manual chart review | Bond screening |
|---|---|---|
| Finding candidates | Coordinator runs an EHR report or scans clinic schedules, then opens charts one at a time | Bond reads every record in scope against the configured criteria; its site cites 10,000+ charts per hour[1] |
| Reading notes | Limited by coordinator time; long histories are hard to read in full | Notes, reports and structured fields are read for every record in scope |
| Time per candidate | The largest share (35.8%) of evaluations took 10 to 30 minutes in one 2012 cancer center study[8] | Coordinator checks the cited evidence; Bond reports 50%+ less chart review[1] |
| Consistency | Two trained coordinators agreed at a kappa of 0.72 on eight calibration charts in one oncology study[9] | The same configured criteria and deterministic scoring rules apply to every record[3] |
| Record of reasoning | Screening log or spreadsheet, with as much detail as time allows | Criterion-to-evidence rationale kept in the audit trail |
| Final eligibility call | Coordinator and investigator | Coordinator and investigator, unchanged |
How is PHI handled during screening?
Screening reads protected health information. Bond signs a business associate agreement with the site.[1] Bond is HIPAA compliant and SOC 2 Type I compliant, and its SOC 2 Type II and ISO 27001 audits are underway.[2]
- Encryption at rest and in transit, AES-256 where applicable[1]
- Role-based access control with SSO support
- Audit logging
- Penetration testing and employee security training
- A Trust Center hosted on Vanta
As of September 2026, the HIPAA Privacy Rule's research provisions include two routes for reviewing records before patient contact: review preparatory to research, where the researcher represents that no PHI will leave the covered entity during the review, and a waiver of authorization approved by an IRB or privacy board.[13] Which route applies, and how it covers an outside vendor's processing, is for the site, its privacy office and its IRB to decide, not Bond. See security and IRB and HIPAA rules for patient outreach.
What does Bond not do?
Identify narrows the search. It does not replace the people who run the study.
- It does not decide eligibility. Bond ranks and explains. The coordinator and investigator make the call, and the screening visit confirms it.
- It does not see what is not in the chart. Care received elsewhere and anything never documented still need a conversation.
- It does not contact patients. Outreach is a separate step in Engage, with scripts the site configures.
- It does not obtain consent. Consent supports the process; the site and PI still obtain consent.
- It does not screen charts without data access. Sites that want to start before the EHR connection is live can begin with list-based outreach.[1]
- It does not remove the need to watch for bias. A 2026 JAMIA study of nine LLMs, using physician-validated patient vignettes, found eligibility judgments largely stable across patient identities. Homelessness produced the largest negative shift, and disparities appeared where a model had to infer behavior or resources.[12] Criteria that turn on adherence or resources are good candidates for human review.
Bring a protocol. We will walk through how Bond would approach its hardest criteria and what your coordinators would see.
Frequently asked questions
Does Identify work with Epic and Oracle Health (Cerner)?
Is Identify priced separately?
Can Identify support feasibility answers?
What happens when the protocol is amended?
When does Bond start reading real patient records?
Sources
- 1.Bond Health: platform overview, FAQ and pricing · Bond Health, 2026
- 2.Bond Health product information · Bond Health, 2026Capabilities, pricing and compliance status described by Bond Health, September 2026.
- 3.Terminology Infrastructure and Graph-Grounded RAG for Clinical Trial Patient Matching · Bond Health, preprint, 2026Goel R. Internal technical report (preprint), August 2026. Not peer reviewed; no public URL yet.
- 4.Retrieval Augmented Generation Enabled Generative Pre-Trained Transformer 4 (GPT-4) Performance for Clinical Trial Screening · medRxiv preprint via PubMed Central (Unlu O et al.); peer-reviewed version in NEJM AI, 2024Quote: "the sensitivity and specificity of determining eligibility for the RECTIFIER was 92.3% (CI) and 93.9% (CI), and study staff was 90.1% (CI) and 83.6% (CI), respectively." Also: "Currently, structured data in the EHR can only be used to determine 5 out of 6 inclusion and 5 out of 17 exclusion criteria." Test set of 1,894 patients; an expert clinician completed a blinded chart review as the reference.
- 5.Manual vs AI-Assisted Prescreening for Trial Eligibility Using Large Language Models: A Randomized Clinical Trial · JAMA (Unlu O et al.), full text via PubMed Central, 2025Quote: "The eligibility rate was 20.4% (458/2242 patients) for the AI-assisted screening method vs 12.7% (284/2234 patients)" and "there were 35 enrollments (1.6%) using the AI-assisted screening method compared with 19 enrollments (0.9%) using the manual screening method". 4,476 patients randomized; enrollment subdistribution hazard ratio 1.79 (95% CI 1.02 to 3.15). Checked against the JAMA full text, September 2026.
- 6.Matching Patients to Clinical Trials with Large Language Models · Nature Communications (Jin Q et al., NIH National Library of Medicine); preprint via PubMed Central, 2024Quote (arXiv preprint 2307.15051 as deposited in PMC): "Manual evaluations on 1,015 patient-criterion pairs show that TrialGPT-Matching achieves an accuracy of 87.3% with faithful explanations, close to the expert performance." Also: "our user study reveals that TrialGPT can reduce the screening time by 42.6% in patient recruitment."
- 7.Cohort selection for clinical trials: n2c2 2018 shared task track 1 · Journal of the American Medical Informatics Association (Stubbs A et al.), 2019Quote: "The best-performing system achieved a micro F1 score of 0.91 using a rule-based approach." Also: "The average kappa score across all criteria was 0.54." Dataset: records of 288 patients, 13 selection criteria, 47 teams.
- 8.Effort Required in Eligibility Screening for Clinical Trials · Journal of Oncology Practice (Penberthy LT et al.), full text via PubMed Central, 2012Single academic cancer center (VCU Massey), 3,467 evaluations over 18 months. Quote: "The largest proportion of evaluations (35.8%) required 10 to 30 minutes, but more than 10% required between 2 to 4 hours for completion." Also: "The average time spent to find, screen, and enroll a patient varied from 3.4 to 8.8 hours".
- 9.Human-AI teaming to improve accuracy and efficiency of eligibility criteria prescreening for oncology trials: a randomized evaluation trial using retrospective electronic health records · Nature Communications (Parikh RB et al.), full text via PubMed Central, 2026355 retrospective charts from a 15-physician community oncology practice in California; University of Pennsylvania IRB. Quote: "Chart-level accuracy, the primary endpoint of Human+AI prescreening is noninferior and superior to Human-alone (76.5% vs. 71.1%). However, efficiency is unchanged with similar average time per chart review, the secondary endpoint, (37.4 vs. 37.8 min)." Calibration on eight charts: "Cohen's Kappa was 0.72 and percent agreement was 86.1%".
- 10.Employing computers for the recruitment into clinical trials: a comprehensive systematic review · Journal of Medical Internet Research, 2014Review of 101 papers on 79 clinical trial recruitment support systems.
- 11.United States Core Data for Interoperability (USCDI) Version 1 · ASTP/ONC Interoperability Standards Platform, 2020Quote: "Clinical Notes • Consultation Note • Discharge Summary Note • History & Physical • Imaging Narrative • Laboratory Report Narrative • Pathology Report Narrative • Procedure Note • Progress Note"
- 12.Sociodemographic bias in large language model clinical trial screening · Journal of the American Medical Informatics Association (Soffer S et al.), 2026Physician-validated vignettes, 33 identity variants, nine LLMs, 58 US Phase II-III protocols. Quote: "Across 58 protocols and 5.3 million evaluations, eligibility judgments were largely stable across identities. ... Homelessness produced the largest negative eligibility shift" and "disparities emerged in domains requiring inference about behavior or resources".
- 13.45 CFR 164.512(i): Uses and disclosures for research purposes · Code of Federal Regulations, via Legal Information Institute (Cornell Law School), 2026Accessed September 2026. Quote: "Reviews preparatory to research. The covered entity obtains from the researcher representations that ... No protected health information is to be removed from the covered entity by the researcher in the course of the review". Waiver route: documentation that an alteration to or waiver of authorization "has been approved by either" an IRB or a privacy board.
Related pages
Engage
Voice and SMS/text agents that pre-screen and schedule the patients Identify finds.
ReadImplementation
What happens between signing and live screening.
ReadIntegrations
How Bond connects to Epic, Oracle Health (Cerner) and the other major EHRs.
ReadBond vs manual chart review
A side-by-side of coordinator time, consistency and audit trail.
ReadUsing the EHR for recruitment
What a site's EHR can and cannot tell you about eligibility.
ReadSecurity
BAAs, encryption, access control and audit logging.
Read