Blog
agentic-aiarchitectureautomationbenchmarksdecision-frameworkevaluationhitlinferencelangfusemodel-analysisopen-sourceproductionragreasoningregulated-industries
July 30, 2026 · 6 min read
Reading the Kimi K3 technical report from the seat of someone who builds agentic systems that have to run on Monday morning: the decisions were made by kernels, caches, and harnesses, not loss curves.
model-analysisagentic-aiinferenceevaluation
May 9, 2026 · 3 min read
Anthropic's agent templates for regulated industries matter less for the customer logos than for the orchestration topology underneath: lazy-loaded skills, governed connectors, and isolated subagents.
model-analysisagentic-aiarchitectureregulated-industries
April 22, 2026 · 7 min read
An honest read of Moonshot's Kimi K2.6 from someone building production agentic systems: why the open weights matter more than the leaderboard, and where it still falls short.
model-analysisagentic-aiopen-sourceevaluation
June 19, 2025 · 4 min read
What two controversial papers, The Illusion of Thinking and its rebuttal, taught us about measuring machine reasoning: many AI failures are benchmark design failures.
evaluationbenchmarksreasoning
March 20, 2024 · 1 min read
When to keep humans in the loop and when to ship fully autonomous agentic systems.
agentic-aihitlautomationdecision-framework
March 15, 2024 · 1 min read
Common ways RAG systems fail in week 2, and how to avoid them with evaluation and observability.
ragproductionevaluationlangfuse