You can ask a language model to extract a sender, request and deadline from an email as JSON.

Getting the values into the exact shape your code expects takes more than asking for a tidy reply. Use a schema where supported, then validate the result.

Why models produce bad JSON

A model generates one token at a time, choosing each from a probability distribution, as covered in our post on training and sampling.

Nothing in that process knows about matching braces. So left to itself a model will, some fraction of the time:

Wrap the JSON in a Markdown code fence or preface it with "Here is the JSON you asked for:".

Add a trailing comma, use single quotes, or leave a comment. It has seen a lot of JavaScript.

Return a number as a string in one response and a number in the next.

Invent a field you did not ask for, or omit one you did, or rename "email" to "emailAddress".

Truncate mid document because it hit the output length limit.

Return an apology instead of JSON when the input is odd.

A one percent failure rate feels fine in testing and is a pager alert at ten thousand requests a day. Treat invalid output as an explicit failure in your application.

Level 1: ask properly

If you are on a model or platform with no structured output feature, prompting gets you most of the way. The pattern that works:

Extract the fields below from the message. Respond with a single JSON object
and nothing else. No explanation, no code fences.

Schema:
{
"sender": string,
"request": string,
"deadline": string in YYYY-MM-DD format or null if none
}

Message:
"""
Hi, this is Dana from accounts. Can you resend the March invoice
by next Friday? Thanks.
"""

Three things are doing the work.

The explicit "nothing else" cuts the chatter. Showing the exact shape, with types, gets consistent field names.

Specify how to represent a missing value, then check whether the model followed that instruction.

Adding one or two worked examples helps further. Set temperature to zero or close to it.

This prompt is a starting point to test.

Measure validity and field accuracy on your own inputs; wording and temperature do not guarantee either.

Level 2: JSON mode

Most hosted APIs and local runners like Ollama have a switch, usually called JSON mode or a response format of json_object, that constrains the model to emit syntactically valid JSON.

Under the hood the sampler is only allowed to pick tokens that keep the output parseable. This eliminates the code fences, the prose and the trailing commas.

It does not enforce your fields.

You can still get valid JSON with the wrong keys or an empty object. Always combine it with the prompt from level 1.

Level 3: structured outputs with a schema

The real fix, and the standard approach in 2026, is to hand the API a JSON Schema and have the model constrained to match it.

The providers call this structured outputs, and it is available from the major hosted APIs and from local runners. The sampler is restricted at each step to tokens that can lead to a document valid against your schema.

Supported schemas constrain the shape, but failures and provider limits still need handling.

The Claude structured output documentation lists refusals and token limits that can interrupt schema compliance.

{
"type": "object",
"properties": {
"sender":   { "type": "string" },
"request":  { "type": "string" },
"deadline": { "type": ["string", "null"], "format": "date" },
"priority": { "type": "string", "enum": ["low", "normal", "urgent"] }
},
"required": ["sender", "request", "deadline", "priority"],
"additionalProperties": false
}

Practical notes:

Mark every field required and set additionalProperties to false. Optional fields lead to the model quietly skipping the hard ones. If a value can be absent, allow null explicitly.

Use enums to describe categorical values, then check the returned value against your allowed list.

Add a description that explains what belongs in a field. Test whether it improves extraction on your inputs.

Give the model room to think when accuracy matters. A schema with a "reasoning" string field first and the answer fields after lets it work through the problem before committing.

Most SDKs let you define the schema as a typed class in your language, Pydantic in Python or Zod in TypeScript, and generate the JSON Schema for you, then parse the response straight into a typed object.

Tool calling returns arguments for a named function. Ordinary tool calling does not guarantee that those arguments match its schema.

A tool named return_result can carry the fields you need, but naming it does not enforce them. On Claude, strict: true enables schema constrained tool inputs for supported models and schemas.

Level 4: validate anyway, and repair

Schema constrained output limits the allowed shape, not the meaning of a value.

A date can be in format but in the wrong year. A string can be empty.

On Claude, unsupported features or excessive schema complexity reject the request with HTTP 400. A refusal or output limit can also leave you without a valid document.

So the code path is always:

response = call_model(prompt, schema)
try:
data = parse_and_validate(response, schema)
except ValidationError as e:
response = call_model(prompt + "\nYour previous reply was invalid: " + str(e)
+ "\nReturn corrected JSON.", schema)
data = parse_and_validate(response, schema)   # give up after this
check_business_rules(data)   # date is in the future, sender is not empty, ...

A bounded retry can address a correctable failure; it can also fail again.

Log the failures. Reading a week of them tells you exactly which prompt sentence to tighten.

Paste a failed response into our JSON formatter to inspect its syntax error. It shows a location when the browser supplies one or reports unexpected end of input; other messages stay unchanged.

Check your schema against the response separately.

Reduce the surface area

Ask for less. A schema with forty fields fails more than four schemas with ten. Split large extractions into passes.

Flat beats deep. Deeply nested structures give the model more places to lose track.

Extract, do not compute. Have the model return the raw quantities and do arithmetic in your code. Models are bad at sums, as the token post explains.

Validate array items separately after parsing. You can isolate item-level validation failures in an already parsed array. A syntax error or truncated response can still prevent the entire array from parsing.

Watch the output limit. If responses are long, raise the max output tokens or paginate. Truncation is the one failure a schema cannot prevent.

Cost and speed

Claude compiles a new schema into a grammar, which adds latency on the first request. Check the provider's limits and caching behavior when comparing response times. JSON itself is verbose, though.

Every quote, brace and key name is tokens you pay for on output.

Shorter keys, compact output and arrays are options to measure. Token counts depend on the text and tokenizer, and a shape change can affect correctness.

The short version

Use structured outputs with a strict JSON Schema, every field required, enums for categories, descriptions on everything.

Keep the prompt explicit about what to return and what to do with missing values. Parse, validate and retry once with the error.

Check business rules in code.

Handle refusals, incomplete replies and business rule failures explicitly before the data reaches the rest of your application.