• Sources: GitHub advisory GHSA-x2rj-828p-hx9m, xinference 2.7.0 release
  • Summary: The Llama3 tool-call parser passes model output to Python eval, so text a request steers the model into producing runs as code in the serving process. The advisory gives the affected range as xinference 2.5.0 and earlier with the fix in 2.7.0, and rates it CVSS v3.1 10.0 because the tested default deployment required no authentication. The path runs from the public /v1/chat/completions endpoint through the tool-call parser rather than through any admin surface.
  • Why it matters: A default xinference deployment exposes that parser on an unauthenticated endpoint, so reaching code execution takes a chat completion request rather than any credential.

send feedback on this story