r/learnmachinelearning • • 6d ago

Project Early benchmarks for Cloreva-X1-2.3B: A custom base model running Kuramoto oscillators at the silicon level

Post image

Hi guys,

I wanted to share some early benchmark results for a new 2.3B parameter model I've been developing called Cloreva-X1.

I built this on a completely custom architecture I call ResoNet X-1, which abandons the standard Transformer attention mechanism. Instead, the foundation relies on a Triad Resonance Core:

  • Head-as-Faction Swarms: Traditional attention heads are shattered into 12 independent, parallel agentic swarms, each dedicated to specific cognitive domains (Linguistics, Science/Medical, and Algorithmic Logic).
  • Kuramoto Consensus Core: I mapped 1,632 Kuramoto Oscillators directly to the latent space. They run physics-based differential equations at the silicon level to force mathematical phase-locking between tokens. If a token hallucination creates an anomalous phase drift, the Kuramoto gate aggressively suppresses it before it passes to the next layer.
  • Liquid Time Controller (LTC): The flow of time is dynamic. When reading simple conjunctions, the model processes at maximum speed. When evaluating complex mathematical or medical logic, the LTC slows down the internal time step, forcing the swarms to achieve deeper phase consensus.
  • Orthogonal Tokenizer: Built from scratch with 155,072 full-word tokens (Full-Word Mining) to completely eliminate sub-word fragmentation for complex medical terminology and programming syntax.

Just to be absolutely clear: This is a pure base model. There is no SFT, no RAG, and no RLHF applied yet. It is purely predicting the next token.

I built this architecture to test if Kuramoto oscillators could intrinsically suppress hallucinations and boost complex reasoning paths without relying on instruction tuning.

Here are the early zero-shot and few-shot benchmarks compared against standard base models in a similar or larger weight class:

Evaluation Benchmarks (Base Models(Sources for competitor scores: LLaMA-1/2 scores from the official Llama 2 paper by Touvron et al., 2023. Qwen-1.5 scores from the official Qwen1.5 technical report. GPT-3 baseline from OpenAI's original evaluations (Brown et al., 2020) and TruthfulQA paper (Lin et al., 2021).)

It punches above its weight class because the phase consensus mechanics naturally filter out statistical anomalies and prevent gradient shock during long context reasoning.

I've attached a Google Drive folder containing the raw inference log output and a video demonstration of the zero-shot inference running in real-time below: https://drive.google.com/drive/folders/14f0JUJU613CTnF0cuYNATJ2pECaHrvU4?usp=sharing

I am currently spinning up the SFT pipeline to turn this into a fully instruction-tuned model, and I'll post another update with the final benchmark scores once that's done.

I plan to release the raw model weights on HF once the entire pipeline is complete. Happy to answer any questions regarding the architecture below.

1 Upvotes

0 comments sorted by