2026-06-29
Top stories
- GLM 5.2 tops Semgrep's IDOR cyber benchmark, beating bare Claude Code Open-weight GLM 5.2 scored 39 percent F1 on IDOR detection versus Claude Code's 32 percent, at 0.17 USD per vulnerability found.
- OpenAI Codex still lacks a way to exclude sensitive files from the model Codex lacks a .codexignore mechanism to exclude secrets and credentials from being read and sent to the model, open since 2025-08-28.
- Developer uses Claude Code to get a second opinion on an MRI report A developer used Claude Code to cross-reference an MRI report against clinical guidelines to question a treatment recommendation.
- Librepods reaches the front page with AirPods reverse-engineered for non-Apple devices Librepods, a GPL-3.0 project, reverse-engineered AirPods control features to expose battery, noise control, and gestures on Android and Linux.
ML research
- Paper argues AI-agent risk should be measured at the repository level A preprint argues that evaluating coding agents one at a time misses ecosystem harm when many agents commit to shared repositories concurrently.
- ToolPrivacyBench audits whether tool-using agents leak private data to the wrong tools 2,150-case benchmark found tool-using agents completed tasks while inappropriately disclosing private data through intermediate tool calls.
Agentic coding
Apple platforms
Infrastructure
Engineering posts
- Cloudflare traces an intermittent image-truncation bug to a discarded Poll::Pending in hyper Cloudflare traced a six-week image-truncation bug to a discarded Poll::Pending in hyper's HTTP/1 state machine, a four-line fix in dispatch.rs.
- HackerRank open-sourced its resume-scoring agent; analysis finds the scores non-deterministic One resume scored 66 to 99 out of 120 through HackerRank's open-source resume-scoring agent, revealing unreliable evaluative judgment.
Markets and companies
Hacker News
- Essay argues age verification is a precursor to speech attribution tops the front page An opinion essay arguing mandatory age verification precedes persistent identity attribution of online speech topped the front page.
- Stanford data page charts memory prices from 1960 to 2026 A Stanford data page plotting per-byte prices for DRAM, flash, and disk memory from 1960 to 2026 reached the front page with 167 points.
- GLM 5.2 benchmark thread, methodology skepticism Hacker News commenters scrutinized the GLM 5.2 benchmark, arguing it pits a bare prompt against Semgrep's scaffolded multi-agent pipeline.