dev.to 25/01/2026 22:12

Taming LLM Output Chaos: A 3-Tier Normalisation Pattern

Leggi la fonte originale
You ask the LLM for "DRAINS" relationships. You get: drains, depletes, exhausts, causes_fatigue, emotionally_draining, negatively_impacts, energy_draining, leads_to_exhaustion, saps_vitality, wears_out, causes_drain... Eleven variations. Per concept. Per run. If your application depends on consistent output from an LLM, you have a problem. I learned this the hard way while building Sentinel, a CLI tool that uses Cognee to detect energy conflicts in personal schedules. This article shares the pattern I developed to normalise chaotic LLM output into predictable, application-ready data. The Problem: LLMs Don't Follow Instructions I was building a knowledge graph where activities could DRAINS energy or REQUIRES focus. Simple enough. I wrote a custom extraction prompt: **REQUIRED RELATIONSHIP TYPES** (use ONLY these exact names): - DRAINS: Activity depletes energy/focus/motivation - REQUIRES: Activity needs energy/focus/resources The LLM nodded along and generated... whatever it felt like. My collision detection algorithm expected DRAINS. Cognee's LLM returned is_emotionally_draining. My BFS traversal found nothing. Tests passed (mocks are liars). Production was broken. The harsh truth: Prompting alone gets you ~70% consistency. For the remaining 30%, you need a normalisation layer. Why Prompting Isn't Enough I tried harder. I added examples. I used few-shot prompting. I YELLED IN CAPS. **CRITICAL**: Use ONLY these relationship types: - DRAINS (not "depletes", not "exhausts", not "causes_fatigue") Result: The LLM now generated DRAINS 85% of the time. But also drains_energy, energy_draining, and my personal favourite: negatively_impacts_emotional_state. The LLM understands semantics, not syntax. It knows these concepts are equivalent. It doesn't care about your string matching. My detection rate: 15-20% of actual collisions found. The Solution: 3-Tier...
Leggi la fonte originale