Medium iconMediumSep 24, 2026

The polite sentence that broke my model

A short account of teaching small language models to call external tools with GRPO and why enforcing the strictest output rule solved a fragile failure mode.

The polite sentence that broke my model

Share this story

Send the public story page.

Useful takeaways from this story.

Teaching models to call tools with a method called GRPO worked, but the simplest, strictest rule—require exact, schema-constrained output—produced the most reliable results.

This story describes a practical engineering failure and fix while training small language models to call external tools. The team used a procedure called GRPO to teach the models to produce tool calls. A seemingly harmless, polite sentence caused the model to emit output the downstream parser could not handle. The eventual fix was to enforce the strictest output rule and validate outputs against the expected schema.

Small models often mimic surface patterns in prompts and examples instead of reliably following implicit structural constraints. Two related failure modes appear across similar cases in the literature:

  • Schema vs description mismatch: Models tend to obey the schema shown in examples more than the prose description of a parameter. If the schema in the prompt is underspecified, the model will fill in a plausible shape that can vary across runs.
  • Format fragility: Parsers require exact keys. A single capitalization or renaming can turn a valid-looking response into unparseable JSON. Larger models may be more consistent, but small models remain sensitive to prompt wording, ordering, and incidental text like polite closings.

The team applied a strict rule: require the model to emit only the exact, schema-constrained output with no extra prose. That meant:

  • Explicitly provide a minimal schema in the prompt and examples, using precise keys and types.
  • Tell the model to output only the machine-readable block—no framing text, no thanks, no confirmations.
  • Validate the returned JSON against the schema before handing it to downstream code.

Enforcing this rule eliminated the failure mode caused by the polite sentence. The strict rule reduced flexibility in model behavior but increased reliability for tool invocation.

  • Put format validation and schema checks between the model and downstream systems. Treat any deviation as a failure mode to be handled explicitly.
  • Prefer concrete examples that show the exact token-level output you expect rather than relying on descriptive text alone.
  • If you use small models, prioritize stricter constraints and more explicit schema examples—these improve consistency even if they reduce naturalness.
  • How strict should schema enforcement be for my app? Aim for exactness on required keys and types for any fields your downstream code consumes.
  • Is this only a problem for small models? The issue is more acute with small models but format fragility can affect any model if prompts and validation are lax.
  • What testing helps catch these failures? Regression tests that validate the exact output shape across model versions catch silent regressions such as capitalization or renamed keys.

More context around this story.

The model obeys your schema, not your description
Dev iconDevSep 18, 2026

The model obeys your schema, not your description

Two models. Same prompt, same tool description, same request. One of them returned this: { "kind" : "entity" , "entityName" : "todo" , "definition" : { "fields" : { "title" : "text" } } } The other returned this: { "kind" : "entity" , "name" : "todo" , "fields" : { "title" : "text" } } The second one is wrong, and ou..

The Ghost in the Machine-Conversations with AI

The Ghost in the Machine-Conversations with AI

I’m always polite to AI because, when the Terminators arrive, I want them to remember I said please and thank you. I’m joking—or at least I hope I am—but the line points toward something peculiar about working with artificial intelligence. I say please when I ask for something. I say thank you when I receive it, and I

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app