1. Evaluation is the product
The hardest part of shipping an agent isn't the prompt or the tools. It's knowing whether the thing is getting better or worse. Teams that ship working agents treat evaluation as a first-class deliverable, not an afterthought.
2. Guardrails are a design problem, not a bolt-on
By the time you're adding content filters and refusal logic to a finished system, it's already too late. Good agents have their constraints baked into the architecture.
3. The data question
If you don't own your own eval data, you don't own your improvement loop. Invest here early.