LLM features fail in two predictable ways: they hallucinate, and they degrade silently. Both have engineering answers. The teams that ship reliably treat the model as an unreliable subsystem and design around it — the way you would treat a flaky third-party API.
Ground every answer in retrievable sources
If a feature cannot point to the document, row, or paragraph that justifies its output, it should not ship. Citations are not a nice-to-have; they are the only way the user can tell a confident answer from a correct one.
Treat prompts as code
- Version them. Diff them. Code-review them.
- Pin model versions. 'Latest' is not a deployment strategy.
- Hold out an evaluation set and run it on every prompt or model change.
- Log inputs, outputs, and the prompt hash. You will need them.
tsconst result = await llm.complete({ prompt: PROMPT_V7, // pinned model: 'gpt-4o-2024-08-06', // pinned inputs, }); await evaluator.record({ promptHash: hash(PROMPT_V7), inputs, output: result.text, citations: result.citations, });
AI features are still software. The discipline that ships reliable software ships reliable AI features.
Md Arifur Rahman is a Senior Software Engineer, Systems Architect, and Cyber Security professional with 8+ years building production-grade platforms across fintech, government, and enterprise SaaS.
