I captured a hundred agent failures a week and fixed three. The bottleneck was curation.
I had full traces and an eval set that was mostly the same bug on repeat. Here is the curation discipline that fixed it.
Search for a command to run...
I had full traces and an eval set that was mostly the same bug on repeat. Here is the curation discipline that fixed it.
My DSPy RAG pipeline compiled clean, passed its hold-out, and still fell apart on live traffic. Here is what I changed.

The eval that finally caught a bad regression for me was not the one in CI. CI had run twelve hours earlier, against a frozen test set, and passed. The regression was sitting in live production traffi
I built a RAG app on LlamaIndex in about four lines. Wire an index to a query engine, point it at my documents, ask a question, get an answer. The first hundred queries were great. I was impressed wit
I built an agent on AWS Bedrock the low-code way. Pick a foundation model, wire an action group to a Lambda, attach a Knowledge Base, add a Guardrail. Then I ran Bedrock's built-in Model Evaluation jo
I tried a bunch of terminal coding agents last month. Some I reopened every morning. Some I uninstalled by Friday. What surprised me is that the ones I dropped were not worse at writing code. The mode
An agent harness is a while-loop, a dispatcher, and a growing message list. The loop you can read in one sitting. The bugs in how you hand tool results back are what actually cost me an afternoon.

Every voice agent I have shipped demoed perfectly. Then a real caller talked over it, changed their mind twice, asked something off-script, and the agent confidently booked the wrong appointment. None

My feed keeps doing this thing. Someone posts that a model helped train the next model, and the top reply calls it the start of an intelligence explosion. Someone else describes a runaway superintelli