agents/prod
07
Observability & evaluation

The gap where quality quietly erodes

Verified  Teams instrument what agents do, but far fewer measure whether output is good. That ~37-point gap is where production quality degrades.

adoption among orgs with production agents
Observability
89%
Evaluation
52%
▲ 37-point gap — LangChain State of Agent Engineering 2025 (1,340 respondents)

Reported  Six leading platforms in 2026: Galileo, LangSmith, Arize, Braintrust, Langfuse, AgentOps. Traditional APM misses hallucinations and semantic drift. Verified LangSmith is framework-agnostic (OpenAI/Anthropic/Vercel/LlamaIndex SDKs) with full OpenTelemetry integration and SDKs in Python, TypeScript, Go, Java.

Debunked
"OpenTelemetry has finalized AI-agent semantic conventions."

Refuted 0–3. The GenAI SIG is actively working on them across CrewAI/AutoGen/LangGraph/Semantic Kernel — but they are not yet finalized. Don't assume a stable spec.