How to Get Reliable JSON Out of an LLM

By Sheng Pang · Published · 7 min read

The most common thing developers ask a language model to do in production is not chat. It is extraction: read this email and give me the sender, the request and the deadline as JSON. Read this receipt and return line items. Classify this ticket and return a category and a confidence. The model is good at understanding the input. Getting it to return exactly the JSON your code expects, every time, is where projects stall. Here is what works.

Why models produce bad JSON

A model generates one token at a time, choosing each from a probability distribution, as covered in our post on training and sampling. Nothing in that process knows about matching braces. So left to itself a model will, some fraction of the time:

  • Wrap the JSON in a Markdown code fence or preface it with "Here is the JSON you asked for:".
  • Add a trailing comma, use single quotes, or leave a comment. It has seen a lot of JavaScript.
  • Return a number as a string in one response and a number in the next.
  • Invent a field you did not ask for, or omit one you did, or rename "email" to "emailAddress".
  • Truncate mid document because it hit the output length limit.
  • Return an apology instead of JSON when the input is odd.

A one percent failure rate feels fine in testing and is a pager alert at ten thousand requests a day. The goal is to make the structure guaranteed, not likely.

Level 1: ask properly

If you are on a model or platform with no structured output feature, prompting gets you most of the way. The pattern that works:

Extract the fields below from the message. Respond with a single JSON object
and nothing else. No explanation, no code fences.

Schema:
{
  "sender": string,
  "request": string,
  "deadline": string in YYYY-MM-DD format or null if none
}

Message:
"""
Hi, this is Dana from accounts. Can you resend the March invoice
by next Friday? Thanks.
"""

Three things are doing the work. The explicit "nothing else" cuts the chatter. Showing the exact shape, with types, gets consistent field names. Saying what to do when a value is missing prevents the model from inventing one. Adding one or two worked examples helps further. Set temperature to zero or close to it. This gets you to perhaps 98 or 99 percent valid output. Not enough on its own, but the foundation for everything below.

Level 2: JSON mode

Most hosted APIs and local runners like Ollama have a switch, usually called JSON mode or a response format of json_object, that constrains the model to emit syntactically valid JSON. Under the hood the sampler is only allowed to pick tokens that keep the output parseable. This eliminates the code fences, the prose and the trailing commas. It does not enforce your fields. You can still get valid JSON with the wrong keys or an empty object. Always combine it with the prompt from level 1.

Level 3: structured outputs with a schema

The real fix, and the standard approach in 2026, is to hand the API a JSON Schema and have the model constrained to match it. The providers call this structured outputs, and it is available from the major hosted APIs and from local runners. The sampler is restricted at each step to tokens that can lead to a document valid against your schema. You get the right keys, the right types and the right enum values, guaranteed.

{
  "type": "object",
  "properties": {
    "sender":   { "type": "string" },
    "request":  { "type": "string" },
    "deadline": { "type": ["string", "null"], "format": "date" },
    "priority": { "type": "string", "enum": ["low", "normal", "urgent"] }
  },
  "required": ["sender", "request", "deadline", "priority"],
  "additionalProperties": false
}

Practical notes:

  • Mark every field required and set additionalProperties to false. Optional fields lead to the model quietly skipping the hard ones. If a value can be absent, allow null explicitly.
  • Use enums for anything categorical. The model cannot invent a fifth category.
  • Add a description to each property. It goes into the prompt and improves accuracy, not just shape.
  • Give the model room to think when accuracy matters. A schema with a "reasoning" string field first and the answer fields after lets it work through the problem before committing.
  • Most SDKs let you define the schema as a typed class in your language, Pydantic in Python or Zod in TypeScript, and generate the JSON Schema for you, then parse the response straight into a typed object.

Tool calling, sometimes called function calling, is the same mechanism with a different framing. You describe a function with a schema for its arguments, and the model's "call" is a schema constrained JSON object. If your API has tool calling but not structured outputs, define one tool called "return_result" and you have the same thing.

Level 4: validate anyway, and repair

Schema constrained output guarantees shape, not sense. A date can be in format but in the wrong year. A string can be empty. And on some platforms the constraint is disabled when the schema is too complex, or a model finishing at its length limit still hands you a truncated document. So the code path is always:

response = call_model(prompt, schema)
try:
    data = parse_and_validate(response, schema)
except ValidationError as e:
    response = call_model(prompt + "\nYour previous reply was invalid: " + str(e)
                          + "\nReturn corrected JSON.", schema)
    data = parse_and_validate(response, schema)   # give up after this
check_business_rules(data)   # date is in the future, sender is not empty, ...

One retry with the error message fixes nearly everything that gets through. Log the failures. Reading a week of them tells you exactly which prompt sentence to tighten. When you are staring at a failed response and cannot see the problem, paste it into our JSON formatter, which reports the line and column of the syntax error, and check your schema against it separately.

Reduce the surface area

  • Ask for less. A schema with forty fields fails more than four schemas with ten. Split large extractions into passes.
  • Flat beats deep. Deeply nested structures give the model more places to lose track.
  • Extract, do not compute. Have the model return the raw quantities and do arithmetic in your code. Models are bad at sums, as the token post explains.
  • Return arrays of items rather than one giant object. If item 37 is broken you lose one item, not the batch.
  • Watch the output limit. If responses are long, raise the max output tokens or paginate. Truncation is the one failure a schema cannot prevent.

Cost and speed

Structured output adds little latency on hosted APIs; the constraint is applied in the sampler. JSON itself is verbose, though. Every quote, brace and key name is tokens you pay for on output. Short key names, no pretty printing in the response, and arrays instead of repeated objects all reduce the bill. If you are sending large JSON to the model as input, minify it first; whitespace is tokens too.

The short version

Use structured outputs with a strict JSON Schema, every field required, enums for categories, descriptions on everything. Keep the prompt explicit about what to return and what to do with missing values. Parse, validate and retry once with the error. Check business rules in code. Do that and JSON from a model becomes as dependable as JSON from any other API, which is to say you still validate it, but you stop being surprised.

← Back to all articles