Lemma Raises $2.3M to Detect Silent AI Agent Failures
AI startup Lemma raised $2.3 million in pre-seed funding to build monitoring infrastructure that detects silent failures where autonomous agents complete tasks but deliver incorrect results.

Lemma, a member of Y Combinator's Fall 2025 batch, secured the $2.3 million pre-seed round from investors including Matrix, Y Combinator, Liquid 2 Ventures, Vermilion Cliffs Ventures, Irregular Expressions, Cervin Ventures, Comma Capital, Position Ventures, and Eight Capital. Angel investors from OpenAI, xAI, Meta, and DoorDash also participated. Founded by Jerry Zhang and Cole Gawin, the startup targets what it calls "semantic failures"—instances where an AI agent executes its workflow without throwing technical errors but still delivers an incorrect outcome, such as citing an outdated policy or hallucinating data.
To address this, Lemma has built an observability layer that has already processed more than one million agent traces. The platform maps out entire execution trees, capturing large language model calls, tool invocations, inputs, outputs, timing data, and retrieval steps. Instead of forcing engineers to manually review thousands of individual logs, Lemma analyzes these production traces against the agent's original instructions, groups recurring issues together, and sends alerts through Slack when critical anomalies arise.
For developers, this infrastructure bridges the gap between identifying a production error and deploying a patch. Lemma integrates directly into development environments using a Model Context Protocol (MCP) server, allowing engineers to query production traces directly from coding tools like Cursor, Claude Desktop, and Claude Code. Once a bug is fixed, the platform converts the real-world failure into an online evaluation test to prevent future regressions. This systematic feedback loop helps practitioners move past unreliable offline evaluations and build more resilient autonomous systems.
This is our own summary of reporting by Unite.AI



