Published onSeptember 12, 2026llmagentsopenaiAnatomy of a Long-Running Agent: Session, Harness, and SandboxDesign long-running agents by separating durable session history, evolving harness logic, and disposable execution environments.
Published onAugust 31, 2026llmagentsopenaiToo Many Tools: Testing Agent Tool DiscoveryEvaluate always-loaded tools, deferred discovery, and programmatic orchestration before expanding an agent tool catalog.
Published onJune 30, 2026agentsopenaievaluationAre Coding Agents Actually Saving Time? Measure Review and ReworkMeasure coding-agent productivity using accepted work, review effort, and later rework instead of generated code or perceived speed.
Published onApril 30, 2026llmopenaiai-engineeringLLM Routing and Fallback: Preserve the Contract, Measure the TradeoffBuild model routing and fallback policies that preserve application contracts while measuring quality, latency, and total cost.