07
Observability & evaluation
The gap where quality quietly erodes
Verified Teams instrument what agents do, but far fewer measure whether output is good. That ~37-point gap is where production quality degrades.
adoption among orgs with production agents
▲ 37-point gap — LangChain State of Agent Engineering 2025 (1,340 respondents)
Reported Six leading platforms in 2026: Galileo, LangSmith, Arize, Braintrust, Langfuse, AgentOps. Traditional APM misses hallucinations and semantic drift. Verified LangSmith is framework-agnostic (OpenAI/Anthropic/Vercel/LlamaIndex SDKs) with full OpenTelemetry integration and SDKs in Python, TypeScript, Go, Java.
Debunked
"OpenTelemetry has finalized AI-agent semantic conventions."
Refuted 0–3. The GenAI SIG is actively working on them across CrewAI/AutoGen/LangGraph/Semantic Kernel — but they are not yet finalized. Don't assume a stable spec.