Harrison Chase

Harrison Chase

@hwchase17 · Twitter ·

I like this framing a lot: agent improvement is harness improvement, not model improvement The interesting interventions are often at the tool boundary: - what context you pass - when tools are provided - how you recover from failure - what gets measured afterward That’s the loop we’re building around deepagents (the orchestration logic) + LangSmith (how you measure)

Arky Yang

Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses — new paper (arXiv, 4 Sep 2026; accepted EMNLP 2026). https://arxiv.org/abs/2609.05736 Treats harness optimization around a fixed model as budgeted selection: edits are guarded intercepts at the tool boundary, not rewrites

Quoted post media