I keep seeing LLM observability come up lately, because once an LLM app goes into production, you find the question is no longer how strong the model is, but whether you can see what it’s doing and where it goes wrong. An LLM isn’t done once you’ve written the prompt; it’s a system that keeps drifting. Without traces, it’s hard to debug why an agent made a particular decision. Without evals, every iteration is just a gamble. Observability is moving toward standards like OpenTelemetry, which means it’s being engineered into something products need. OpenAI’s Agent Builder, launched two months ago, fills in this piece for people behind the scenes too. Traditional observability giants like Datadog have also made LLM observability a product line, with special emphasis on monitoring and experimentation workflows for agentic systems. The direction is already very clear. What’s more interesting is that while all these commercial tools are being discussed, open-source, self-hosting-friendly options like Langfuse are easy to skip over. Its README makes clear that it wants to fill in the whole skeleton that “turns LLM engineering into an iterative process”: tracing, metrics, prompt management, evaluation, datasets and experiments. And its integrations are genuinely practical. I think there’s a sharp contrast behind this. Right now everyone wants speed, chasing launches, chasing packaging, chasing the rhythm of a demo that “looks like it works.” But the teams that make products last and stay stable all end up back at the same thing: can you leave a record of the system’s behavior, quantify it, and turn failures into test assets for the next round of iteration? That’s why 2025–2026 has seen a wave of “LLM observability tools compared” articles (though not many in the Traditional Chinese world). People are filling in a new era’s fundamentals. It resonates with me because it feels a lot like building knowledge infrastructure. If you don’t keep the context and connect the chain of evidence, you can only iterate on memory and emotion. LLMs are the same: if logs, traces and evals don’t become routine, a product easily gets stuck in the loop of “how did it break again?” In short, if you’re building RAG, agents or any LLM feature that touches production, I recommend looking at Langfuse first, at least to get the skeleton of trace + eval + prompt management up quickly. I’m also looking forward to future versions covering OpenTelemetry’s GenAI semantic conventions more smoothly, so LLM observability really becomes a standard every team can plug into.
AI Observability Is Becoming Standard
Translated from the Chinese original · Read the original
LocatorVibe Coding Valley
2 wks
Overview
End of the trail
415 words, and you made it to the end.
Newsletter
Get the next essay by email.
One email when a new essay goes up, nothing else. Unsubscribe in one click.
← Previous · 2025.12.05How Google Came Back Next · 2025.12.09 →A Real Botnet Encounter: The React Vulnerability Rated CVSS 10
Nearby cairns · the closest essays on the map
Vibe Coding Valley3 days apart How Google Came Back Google looked stuck in the innovator's dilemma in 2022–2023, then caught up and briefly pulled ahead in 2025. Notes on a WSJ piece about the three things it got right. Vibe Coding Valley12 days apart Professional Software Developers Don't Vibe, They Control A study of 99 professional developers using AI agents for coding in 2025. It reflects the state of things before GPT-5.2 and Claude 4.5, but it's still well worth a look. Vibe Coding Valley1 day apart A Real Botnet Encounter: The React Vulnerability Rated CVSS 10 My server was quietly running someone else's malware for eleven hours. How I found it in journalctl, traced it to CVE-2025-66478 in Next.js, and what we're changing so it doesn't happen again.