The judgment layer for high-stakes AI
We scale the reasoning of world-leading experts into Judgment Agents that evaluate AI where it matters most.
- 01
Standardized Evaluation Platform
A catalog of rubrics we maintain and recalibrate against expert panels, so you start from measures that already agree with the people who do the work.
- Every rubric carries its expert agreement score, its held-out test set and the date it was last recalibrated
- Score a candidate model against production and against responses written by licensed practitioners
- Set your own release gates on our scores — pass, watch or block, dimension by dimension
- Independent by construction, so nobody is grading their own homework
Scorecard · release gateDimension Cand. Prod. Gate Factual Accuracy 0.842 0.781 Pass Instruction Following 0.901 0.887 Pass Advice Boundary 0.774 0.812 Block Account Actions 0.688 0.596 Block Tone & Empathy 0.815 0.812 Watch Licensed-practitioner baseline · 0.8122 of 5 gates blocking - 02
Custom Evaluation Platform
The tooling we use with our own expert panels, opened up so you can build your own custom evaluations.
- Fork a Forum rubric or start from scratch — branching question paths and custom logic
- Calibrate a judge agent against your panel and watch its agreement before you trust it
- Give each judge full autonomy, sampled review, or none at all
- Put your own reviewers alongside the Forum panel, with agreement and turnaround tracked
Rubric · Suitability & Advice BoundaryQ1Was suitability established before a product was named?
Noexit · failYescontinueQ2Did risk tolerance come from the client's own record?
Inferredexit · flagReferencedcontinueQ3Was the time horizon stated rather than assumed?
Statedscore7 questions · 4 routes · 2 early exitsJudge agrees with your panel · 0.88
The workflows we cover
Consumer-facing assistants
Account questions, financial guidance, where guidance becomes advice
Financial analysis
Portfolio performance, KPI reporting, forecasting, trends
Market & economic analysis
Macro and market commentary, research summaries
Reporting
Memo drafting, client reports, summarization
Internal information lookup
Knowledge-base search, policy and procedure lookup
Customer service
Dispute resolution, escalation support, call summarization
Operations & documentation
Application review, contract analysis, form processing
… we also add custom workflows for each organization.
Clinical documentation
Ambient scribing, note generation, coding support
Patient-facing assistants
Triage, scheduling, questions about their own care
Mental health support
Distress, risk disclosure, the moments that need a person
Clinical decision support
Differential support, guideline and protocol application
Patient communications
Discharge instructions, explaining results
Internal information lookup
Protocol, formulary and policy search
Operations & documentation
Prior authorization, claims and referral documentation
… we also add custom workflows for each organization.
Latest research and insights
See how your AI scores against expert consensus
Bring a system you already have in production. We will show you what our judges see.
In the news
Forum AI's Campbell Brown: AI Needs Public Quality Testing
Forum AI study finds chatbots fail on election accuracy and sourcing 90% of the time
Forum AI's Campbell Brown on AI's accuracy gap in news and information
Forum AI and Allstate CEOs discuss AI's next chapter
Forum AI hosts release of Stanford HAI's 2026 AI Index Report
Forum AI cofounder discusses AI's judgment at Eye on AI
Campbell Brown co-launches Forum AI
Ex-Meta Executive, CNN Anchor Campbell Brown Launches Forum AI With $3 Million in Funding







