Taming LLM Output Chaos: A 3-Tier Normalisation Pattern
Leggi la fonte originale
You ask the LLM for "DRAINS" relationships.
You get:
drains, depletes, exhausts, causes_fatigue,
emotionally_draining, negatively_impacts,
energy_draining, leads_to_exhaustion,
saps_vitality, wears_out, causes_drain...
Eleven variations. Per concept. Per run.
If your application depends on consistent output from an LLM, you have a problem. I learned this the hard way while building Sentinel, a CLI tool that uses Cognee to detect energy conflicts in personal schedules.
This article shares the pattern I developed to normalise chaotic LLM output into predictable, application-ready data.
The Problem: LLMs Don't Follow Instructions
I was building a knowledge graph where activities could DRAINS energy or REQUIRES focus. Simple enough. I wrote a custom extraction prompt:
**REQUIRED RELATIONSHIP TYPES** (use ONLY these exact names):
- DRAINS: Activity depletes energy/focus/motivation
- REQUIRES: Activity needs energy/focus/resources
The LLM nodded along and generated... whatever it felt like.
My collision detection algorithm expected DRAINS. Cognee's LLM returned is_emotionally_draining. My BFS traversal found nothing. Tests passed (mocks are liars). Production was broken.
The harsh truth: Prompting alone gets you ~70% consistency. For the remaining 30%, you need a normalisation layer.
Why Prompting Isn't Enough
I tried harder. I added examples. I used few-shot prompting. I YELLED IN CAPS.
**CRITICAL**: Use ONLY these relationship types:
- DRAINS (not "depletes", not "exhausts", not "causes_fatigue")
Result: The LLM now generated DRAINS 85% of the time. But also drains_energy, energy_draining, and my personal favourite: negatively_impacts_emotional_state.
The LLM understands semantics, not syntax. It knows these concepts are equivalent. It doesn't care about your string matching.
My detection rate: 15-20% of actual collisions found.
The Solution: 3-Tier...
Leggi la fonte originale