- Sources: arXiv 2608.24571
- Summary: The paper presents SMITH, for Schema-grounded Multi-task Iterative Tool Honing, which trains tool creation and tool use in a single policy rather than prompting a frozen model at inference time. Training uses separate reward axes for schema, code, and outcome. The authors report a 4B Qwen3 reaching 79.8 macro-average accuracy against an untrained 30B-A3B tool-writer.
- Why it matters: Existing tool-creation systems never tell the writer that the schemas it emits are schemas it will have to invoke, and the reported gains come from closing that loop.
send feedback on this story