LLM Returns Invalid JSON — Diagnosis & Fix Guide
Why an LLM that returns JSON in testing starts returning markdown fences, prose preambles, or truncated objects in production — and the four fixes, in the order worth trying them.
You asked the model for JSON. It gave you JSON for two weeks. Now JSON.parse is throwing in production, on maybe one request in fifty, and the logs show something like this:
SyntaxError: Unexpected token ` in JSON at position 0This guide covers the four distinct causes behind that error and the order to address them in.
🔍 First: find out which failure you have
Before changing anything, log the raw response body on parse failure. The four causes look identical from the exception and completely different from the payload.
try {
return JSON.parse(raw);
} catch (err) {
// Log the payload, bounded. The exception alone tells you nothing useful.
console.error('LLM JSON parse failed', { raw: raw.slice(0, 500) });
throw err;
}| What the payload looks like | Cause | Section |
|---|---|---|
```json { ... } ``` |
Markdown fencing | 1 |
Sure! Here is the JSON: {...} |
Prose preamble | 1 |
{"items": [{"id": 1}, — stops dead |
Truncation | 2 |
| Valid JSON, wrong or missing fields | Schema drift | 3 |
1. 🛠️ Fencing and preambles — use constrained decoding
Chat-tuned models are trained to present code nicely to humans. Asking for "only JSON, no markdown" lowers the rate; it does not eliminate it, and the rate changes every time the provider updates the model behind the same name.
❌ Prompting and hoping
const prompt = 'Return ONLY valid JSON. No markdown. No explanation.';
const raw = await model.generate(prompt + input);
return JSON.parse(raw); // throws on the polite responses✅ Constraining the decoder
Every major provider now supports forcing the output to a schema. The decoder is physically prevented from emitting a token that would make the output invalid, so fences and preambles cannot occur:
const result = await ai.run(MODEL, {
messages,
response_format: {
type: 'json_schema',
json_schema: {
type: 'object',
properties: {
category: { type: 'string', enum: ['billing', 'bug', 'feature', 'other'] },
estimatedHours: { type: 'number' },
confidence: { type: 'string', enum: ['low', 'medium', 'high'] },
},
required: ['category', 'estimatedHours', 'confidence'],
},
},
});Tip
Use enum wherever the value belongs to a fixed set. A free-form string field will eventually return "Billing " with a trailing space or "bug/feature", and the switch statement consuming it will fall through to a default nobody tested.
If your provider genuinely has no structured-output mode, then — and only then — strip defensively, and treat it as a temporary measure:
// Fallback only. This handles the failure you have seen, not the next one.
const cleaned = raw.replace(/^[^{[]*/, '').replace(/[^}\]]*$/, '');2. 🛠️ Truncation — check the finish reason
If the payload stops mid-object, the JSON is not malformed. It is incomplete, because generation hit the output token ceiling.
if (response.finishReason === 'length') {
// Not a parse problem. The model was cut off.
}Three fixes, best first:
- Ask for less. Return ids instead of full records; return the top five, not everything. Most truncation is a schema asking the model to restate data you already have.
- Raise the output limit. Cheap, but it moves the ceiling rather than removing it.
- Paginate the work. For genuinely large outputs, loop over chunks rather than requesting one enormous object.
Warning
Truncation scales with input length as well as output — longer inputs produce longer outputs. A feature that is stable on typical data will fail on your largest customer, which is the worst possible group to discover it with.
3. 🛠️ Schema drift — validate, do not trust
Structurally valid JSON with the wrong shape is the most dangerous case, because nothing throws. The field is simply undefined three functions away.
import { z } from 'zod';
const ClassificationSchema = z.object({
category: z.enum(['billing', 'bug', 'feature', 'other']),
estimatedHours: z.number().positive().max(200),
confidence: z.enum(['low', 'medium', 'high']),
});
// Constrained decoding shapes the output; this proves it.
const parsed = ClassificationSchema.safeParse(result.response);
if (!parsed.success) {
// A caught, logged, handleable failure — not a silent undefined downstream.
return fallbackClassification();
}Note the .max(200) on hours. Constrained decoding guarantees a number; it does not guarantee a sensible number. Range checks on numeric fields catch the responses that are valid, parseable, and absurd.
4. 🛠️ Decide the fallback before you need it
Every LLM call needs an answer to "what does the user see when this fails?" — and for a parse failure specifically, retrying the identical request is rarely it. If the model produced unusable output once, the same prompt at the same temperature will often do it again.
Better options, in order:
- Serve a degraded result. The AI-assisted part is absent; the feature still works. Correct for suggestions, summaries, and classifications with a sane default.
- Retry once, with a lower temperature. Worth one attempt, not three.
- Queue for human review. Correct when the output feeds something consequential.
Whatever you pick, put the failure on a counter. A parse failure rate that quietly rises from 1% to 20% after a provider-side model update is invisible unless something is watching it.
✅ Checklist
[ ] Raw payload logged (bounded) on every parse failure
[ ] response_format json_schema used where the provider supports it
[ ] enum constraints on every fixed-set string field
[ ] Finish reason checked for `length` before blaming the parser
[ ] Runtime validation with range checks, not just a cast
[ ] A defined, tested fallback path
[ ] Failure rate on a counter with an alertWork through it in that order. The first two items remove most of the volume; the rest is what stops the remainder becoming an incident.
Frequently asked questions
Most chat-tuned models are trained to format code for human readers, so a request for JSON produces a fenced block. Prompting against it reduces the rate but never to zero. The reliable fix is constrained decoding via a response_format json_schema parameter, which forces the decoder to emit only tokens valid for your schema.
Only as a last-resort fallback for a provider with no structured-output support. A regex handles the failure you have seen and not the next one — a prose preamble, a trailing explanation, or two JSON objects in one reply. Constrain the output first and keep any stripping as a narrow safety net.
Usually truncation. The response hits the max output token limit mid-object, so the JSON is genuinely incomplete rather than malformed. Check the finish reason on the response — if it is length rather than stop, raise the output limit or reduce what you are asking the model to return.
Yes. Constrained decoding makes structurally invalid output rare, not impossible, and says nothing about whether the values make sense. Validate with Zod or an equivalent so a bad response becomes a caught, logged exception rather than an undefined field somewhere downstream.
Experiencing a similar issue?
Describe it and get an automated scoping estimate in seconds, with the option to book a free 30-minute diagnostic call.