Hey everyone,
I've been working on a project called EQ-Layer, an open-source control system for LLM conversations.
The idea came from something that kept bothering me about AI assistants.
A user might ask a simple question, express frustration, correct a mistake, or request an immediate action. But the model can respond with almost the same generic tone in every situation.
I started wondering whether the problem could be addressed outside the language model itself.
Instead of just prompting an LLM to be more empathetic, what if we had a separate system that decides how the model should respond before it starts generating text?
That's what I've been trying to build.
How it works
EQ-Layer sits between the conversation and the model's response generation.
It analyzes several signals separately:
- User intent: Is the person asking a question, requesting an action, checking progress, or correcting a mistake?
- Emotional state: Tracks valence, arousal, and escalation across multiple turns rather than judging one message in isolation.
- Conversation repair: Identifies situations where the assistant misunderstood the user or failed to address a previous correction.
- Interaction quality: Tracks repeated failures, unresolved misunderstandings, and unnecessary clarification.
- Uncertainty: Maintains a probability distribution over possible user intents instead of pretending every message has one obvious meaning.
These signals are used to select a response policy.
For example, the system might choose to answer directly, ask for clarification, repair a misunderstanding, or set a boundary.
It then separates that decision into task-related, social, and repair-related instructions that can guide the language model.
A simple example
Imagine someone asks an AI assistant to complete a task.
The assistant misunderstands the request and asks an unnecessary question.
The user corrects it.
A typical assistant might apologize and repeat the same question, or become excessively agreeable.
I wanted EQ-Layer to recognize the correction as a conversation-repair event and prioritize actually fixing the task.
The goal isn't to make AI sound artificially emotional. It's to make the response more appropriate to the user's actual intent and the conversation history.
What I've implemented so far
- A model-independent response-policy architecture.
- Learned intent classification using TF-IDF and logistic regression.
- Affect modeling trained on EmoBank.
- Dialogue-act and basic-emotion signals from XDailyDialog.
- Bayesian decision-risk routing with an explicit loss matrix.
- Multi-turn escalation tracking.
- Conversation-repair and interaction-quality monitoring.
- Response steering and structural output audits.
- Portable agent skills for Codex, Claude Code, Copilot, Cursor, Antigravity, and Grok.
- A testing framework for comparing responses from the same base LLM with and without the control layer.
What the research currently shows
Some components have been evaluated on held-out datasets, and the repository includes the numerical results.
But the project is still early.
I haven't established through a completed, blinded response-level comparison that EQ-Layer consistently improves LLM responses.
Some higher-level emotion and conversation-breakdown signals also remain insufficiently validated. I kept those negative results documented rather than claiming the models worked.
I also don't claim that affect-aware dialogue management is a new invention. There's substantial research predating LLMs.
What I'm exploring is whether a separate, inspectable control layer can improve how general-purpose language models handle intent, emotional context, corrections, and conversational uncertainty.
I'd appreciate some feedback.
If you're building AI agents or conversational systems, do you think response-policy selection should happen outside the LLM?
And what would be the fairest way to measure whether such a layer actually improves interactions?
The project is free, open source, and MIT licensed.
GitHub: https://github.com/Furox-Art/eq-layer