Adding an AI Feature to a Product That Already Works
Most AI integration work in 2026 is not model training — it is wiring a model into a system that already has users, a bill, and an uptime expectation. The engineering that makes that survivable.
The most common request I get on the build side is no longer "build me an app." It is "we have an app, add AI to it." Upwork's own data puts AI integration among the fastest-growing freelance categories, and the role clients describe when they say "AI developer" is almost never a research role — the model is rarely the hard part.
The hard part is that you are attaching a slow, non-deterministic, metered dependency to a codebase that currently assumes none of those things. Here is what that actually takes.
The model is a network call that can lie
Treat it as a third-party API with three unusual properties, and most of the design follows:
- It is slow. Seconds, not milliseconds, and the tail is much worse than the median.
- It is metered. Every call has a price, and the price scales with input you may not control.
- It can be confidently wrong. Not "returns a 500" wrong — "returns a well-formed answer that is false" wrong.
Every failure I have been called in to fix traces back to a codebase that handled the first two and ignored the third.
Never parse prose
The single highest-leverage decision in an LLM integration is refusing to accept free-form text where you need data.
// ❌ This works in the demo and fails in production.
const reply = await model.generate(`Classify this ticket. Answer with one word.`);
const category = reply.trim().toLowerCase();That code is one polite preamble away from category === "sure! the category is billing". Regexing your way out of it is a losing game — you are writing a parser for an output space the model redefines with every prompt tweak.
Use constrained decoding instead. Every serious provider now supports forcing output to a JSON schema:
const result = await ai.run(MODEL, {
messages,
response_format: {
type: 'json_schema',
json_schema: {
type: 'object',
properties: {
category: { type: 'string', enum: ['billing', 'bug', 'feature', 'other'] },
confidence: { type: 'string', enum: ['low', 'medium', 'high'] },
},
required: ['category', 'confidence'],
},
},
});
// Then validate anyway — the schema constrains the decoder, not your type system.
const parsed = ClassificationSchema.parse(result.response);Important
Constrain and validate. Constrained decoding makes malformed output rare; it does not make it impossible, and it says nothing about whether the values are sane. The Zod parse is what turns "rare weird failure at 2am" into "a caught exception with a stack trace."
The enum matters as much as the schema. A free-string category will eventually return "Billing " with a trailing space, or "billing/account", and your switch statement will fall through to a default that nobody tested.
Decide what happens when it is down
Your product had an uptime number before you added this. Ask, for each new AI call: if this takes twelve seconds, or fails, what does the user see?
There are only three honest answers, and you should pick one deliberately per feature:
- Degrade. The AI part disappears and the rest of the page works. Correct for summaries, suggestions, autocomplete — anything additive.
- Queue. The request is accepted, the work happens out of band, the user is told. Correct for anything expensive or long.
- Fail loudly. The operation genuinely cannot proceed. Correct for very little, and if it is your answer for a core flow, reconsider whether the model belongs on the critical path at all.
On an edge runtime this is often structural rather than a try/catch. Work that should not block the response goes to a post-response hook — on Cloudflare Workers, ctx.waitUntil() — so the user gets their page and the model call finishes on its own time.
Know the bill before you get it
Token cost is usage-driven, and usage is the thing you cannot predict from a demo. Two habits prevent the surprise:
Cap the input. Truncate user-supplied text before it reaches the prompt. An unbounded field is an unbounded invoice, and a user who pastes a 400-page PDF into your support box is not malicious, just ordinary.
Rate limit the open paths. Any unauthenticated endpoint that reaches a model needs a limit keyed to something real — IP, session, account. This is not abuse prevention theatre; it is the difference between a bad day and a bad month.
// Every open POST on this site that touches a model or a paid API
// passes through a KV-backed limiter before it gets anywhere near the model.
const limited = await rateLimit(env, `classify:${ip}`, { limit: 10, windowSeconds: 3600 });
if (limited) return tooManyRequests(limited.retryAfterSeconds);Build the evaluation set on day one
This is the step teams skip, and it is the one that decides whether the feature can be improved after launch.
An evaluation set is boring: twenty to fifty real inputs with the output you would accept, in a file, run by a script. That is all. Without it, every prompt change is a vibe, and "the new prompt seems better" is a claim nobody can check — including you, three months later, when you have forgotten which cases the old prompt handled.
Tip
Harvest the set from production. Log the inputs to your AI feature from the first day, review a sample weekly, and add every case that surprised you. Within a month you have a regression suite that reflects your actual users rather than your imagination of them.
What this means for scoping
If you are hiring for this, the useful questions are not about models. They are:
- What does this feature do when the provider is slow or down?
- Where is the output validated, and what happens to output that fails validation?
- What is the per-request cost ceiling, and what enforces it?
- How will we know a prompt change made things better?
A proposal that answers those is describing production. One that answers only "which model" is describing a demo.
Related
AI & agent integration is one of the build packages here, and this guide covers the specific failure — a model returning unparseable JSON — in more depth. If you are weighing whether a model belongs in your product at all, the free intro call is a reasonable place to test that out loud.
Hitting something like this in your own codebase?
Describe the problem and get an automated scoping estimate in seconds, with the option to book a free 30-minute diagnostic call.