- Sources: preprint
- Summary: An unreviewed preprint evaluates exposing tools as typed Python stubs invoked through code, with execution and results handled in one agent turn. It reports a 10.6 percent gain over the JSON baseline for the GPT-5.6 family and parity or better in 11 of 14 models on BFCL v4. Under context rot the programmatic form held stable while the JSON baseline degraded 2.3 percent.
- Why it matters: The result is a concrete interface choice for anyone building an agent harness, pending independent replication.
send feedback on this story