/ About this event
AI is becoming powerful enough to act. Now it needs to become trustworthy enough to deserve that power.
Let’s create a sandbox for AI Trust together.
As LLMs and AI agents move from generating text to influencing real decisions, the cost of confident mistakes gets much higher. A polished answer isn’t enough. We need systems that can verify claims, evaluate evidence, surface uncertainty, detect fragile reasoning, and decide when an output is trustworthy enough to act on.
Decision-Grade AI is a Stanford-based panel discussion + buildathon bringing together technical builders to prototype that missing layer.
Over the course of the event, we will have an exclusive panel discussion and buildathon.
Opening Talk: Kiana Jafari, Executive Director @ Stanford Center for AI Safety
Featured Panelists: Bryant McCombs, GTM @ OpenAI Clark Barrett, Computer Scientist @ Stanford Wesley Deng, Senior Researcher @ Microsoft Sarath Swaminathan, Research Scientist @ IBM
What You’ll Build:
AI is getting better at producing answers. We want to build the systems that determine whether those answers deserve to drive decisions.
Teams will tackle one of five core challenges -
→ Claim verification - break AI outputs into testable claims and check them against evidence → Citation integrity - determine whether sources actually support what a model says → Trust scoring - create signals and interfaces that communicate reliability and uncertainty → Robustness + bias - test how answers change under different framing, assumptions, or prompts → Safer agents - design monitoring and verification systems before AI agents take consequential actions
We’ll start with a technical kickoff introducing the problem space, example architectures, tools, and build directions. After the expert panel, teams will choose a challenge and build with mentor support before presenting working prototypes in final demos.
This is not a hackathon about building another chatbot.
It’s about building the infrastructure underneath the next generation of AI systems - the tools that help answer:
Can we verify this? What evidence supports it? How confident should we be? Should a human or agent actually act on it?
Our team has been exploring these questions through Ezio, where we’re developing approaches to evaluating AI outputs in real time. This buildathon opens the challenge up to a broader community of builders interested in making AI systems more reliable, auditable, and decision-ready.
Who should attend:
CS, engineering, and data science students; AI/ML researchers; software engineers; technical founders; product builders; recent technical alumni; and designers interested in LLM evaluation, verification, model robustness, or agent reliability.
You don’t need prior experience in AI safety.
You just need to want to build systems that make AI harder to fool - and safer to trust.
Exceptional builders. One day. Build the trust layer for AI.
SCHEDULE:
10:30–11:00 AM | Check-in, coffee & team mixing 11:00 AM–12:00 PM | Technical kickoff & demos 12:00–1:00 PM | Expert Panel Discussion 1:00–1:15 PM | Challenge reveal & team formation 1:15–3:00 PM | Build with mentor support 3:00–3:30 PM | Final demos, judging & awards
Cohost 1 - TokenRouter TokenRouter is a verified model gateway built for developer and enterprise teams — launched under PaleBlueDot AI (Series B, B Capital–backed).
Cohost 2 - Ezio Ezio is the trust layer for AI agents and LLM outputs, measuring AI-generated responses before they are used in high-stakes workflows. This event is a part of #SFTechWeek—a week of events hosted by VCs and startups to bring together the tech ecosystem. Learn more at www.tech-week.com

