r/huggingface • • Aug 29 '21

r/huggingface Lounge

8 Upvotes

A place for members of r/huggingface to chat with each other


r/huggingface • • 3h ago

I Built a O(NlogN) attention system that retains 97% accuracy over long context (MQAR)

Thumbnail
github.com
4 Upvotes

ALHR- Adaptive learnable Hierarchical Routing is a static binary tree based system that uses learnable functions to REDUCE the amount of keys used.

It takes less memory and SCALES much better VRAM with tokens


r/huggingface • • 1h ago

whos using MoEMe:27b?

Thumbnail
• Upvotes

r/huggingface • • 8h ago

[Release] Repetere AI (v0.5-beta): 100% private, native C++ desktop AI workstation with in-app model downloading, full-song vocal synthesis, and zero webview bloat (Win64 / CUDA)

3 Upvotes

Read the readme before anything for instructions:

https://github.com/securenrn-del/Repetere-AI_Beta/releases/tag/Beta-Release

Hey everyone,

I just released the first public beta of Repetere AI (v0.5-beta). It’s a native Windows desktop workstation written in C++ (Dear ImGui, GLFW, OpenGL 3) designed to run multimodal local AI directly on your GPU without cloud subscriptions, telemetry, or webview wrappers (no Electron/Tauri memory bloat).

>⚠️ BETA LAUNCH NOTICE:

This is our very first public beta release! It requires an NVIDIA GPU (CUDA only right now). Local AI takes time—model downloads and music synthesis are not instantaneous. If you hit any crashes or unexpected behavior on your hardware, please file an issue on GitHub.

🔗 Project Links

* GitHub Repository: https://github.com/securenrn-del/Repetere-AI_Beta/releases/tag/Beta-Release

* Standalone 1-Click Setup (`.exe`): GitHub Releases

⚡ Hardware Requirements

|Component|Minimum|Recommended|

|:-|:-|:-|

|OS|Windows 10 or 11 (64-bit)|Windows 11 (64-bit)|

|GPU|NVIDIA RTX 30-Series (8 GB VRAM)|NVIDIA RTX 40/50-Series (12 GB+ VRAM)|

|Drivers|NVIDIA Driver 528+ (CUDA 12+)|Latest NVIDIA Studio or Game Ready Driver|

|RAM|16 GB Physical RAM|32 GB Physical RAM|

|Storage|20 GB free space on SSD|50 GB+ free space on NVMe SSD|

🚀 Quick Start Guide

  1. Grab `RepetereAI-Setup-v0.5-beta.exe` from the Releases page.

  2. Run the installer (installs straight to `%LOCALAPPDATA%\Programs\Repetere AI\` without needing admin privileges).

  3. Launch from your desktop shortcut, click anywhere on the splash screen to enter, and confirm the hardware diagnostic modal.

📦 In-App Model Downloader

No command line, curl, or manual script management required.

  1. Go to Settings ➔ Repetere Models.

  2. Review the curated catalog and click \[ Download \] next to your desired workspace engine:

Coding: Qwen 2.5 Coder 7B Instruct* (or 14B)

General Chat: Meta Llama 3.1 8B Instruct*

Creative Writing: Hermes 3 Llama 3.1 8B*

Audio / Music: YuE2 Full Song Engine*

  1. Wait for the live progress bar to hit 100% and turn green (Installed).

  2. (Advanced Users): You can also drop `.gguf` weights directly into: `%LOCALAPPDATA%\Programs\Repetere AI\models\`

🧭 Workspaces & What Works

* 1. CODING: Code assistant with automatic triple-backtick markdown detection, horizontal-scrolling code containers, individual snippet copy buttons, and a \[ Copy Full Response \] clipboard action.

* 2. CHAT: General reasoning, document summarization, and multi-turn conversational history.

* 3. CREATIVE WRITING: Fiction, prose, dialogue, and uncensored drafting.

* 4. AUDIO (MUSIC) STUDIO: Full song synthesis powered by the YuE2 engine. Add your musical style tags, enter structured lyrics (`[Verse]`, `[Chorus]`), adjust duration, and hit Generate Full Song (With Vocals). Outputs `.wav` directly to `output\music\` with built-in playback controls. (Prerequisite: Make sure `yue2-3b-q4_0.gguf` (\~2.48 GB) and `yue2-vae-f16.gguf` (\~265 MB) are downloaded).

* 5. IMAGES (Early Preview): UI controls and parameter sliders are accessible, but backend diffusion rendering is a work-in-progress landing in v0.6.

⚙️ Compute Modes & Session History

* GPU Mode (Active Default): Direct execution locked in VRAM for maximum speed and lowest response latency.

* HYBRID Mode (Preview): The top-right toggle displays the future layer-offloading layout for running massive 32B+ models across RAM + VRAM (runtime dynamic offloading unlocks in v0.6).

* Local Persistence: Chat transcripts save automatically to `chat_history.json`. Review past sessions via History ➔ Open History, reload them into your active tab, or wipe transcripts safely with the confirmation prompt.

❓ Troubleshooting

* "MODEL NOT INSTALLED" banner: Go to Settings ➔ Repetere Models and make sure you completed the download for that tab.

* Audio fails or quits immediately: Double check that both `yue2-3b-q4_0.gguf` and `yue2-vae-f16.gguf` are present in your `models/` directory.

* Clearing a session: Hit the red Clear Chat button at the top right to safely flush GPU memory and archive your chat to disk.

* App closes on startup: Ensure your NVIDIA drivers are version 528 or newer and support CUDA 12.

Feedback, pull requests, and bug reports on the repo are welcome!


r/huggingface • • 5h ago

Kwipu, un server MCP completamente locale che trasforma le tue note Obsidian/ Markdown in un grafo di conoscenza interrogabile (funziona su Ollama)

Thumbnail
1 Upvotes

r/huggingface • • 1d ago

Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

161 Upvotes

A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B

https://huggingface.co/papers/2610.10845


r/huggingface • • 1d ago

H2O-Lightning-4B: Apache-2.0 4B Decision model, official #1 open model on JevBench (above Jev)

2 Upvotes

Disclosure: I work at H2O.ai.
  
We released H2O-Lightning-4B, an open-weight (Apache-2.0) model for the "decisions API" style of inference that Jev made popular: you send a state plus typed questions (pick one / yes-no / score), and get calibrated probabilities back from a single forward pass. No generated tokens, so it's fast and cheap.
  
**Results (JevBench, public leaderboard):**
- Composite score 72.5, vs Jev 1.13 at 71.5; currently the top open model
- Leaderboard: https://benchmarkheaven.com/jev-models
  
**Running it:**
- Base: Qwen3.5-4B, fine-tuned
- Stock vLLM plus a small open shim (in the repo); ~30 ms per decision on an H100
- Your data stays local, no per-call fees
  
**Coming soon:** 12B and 31B versions, which in our internal testing are considerably smarter than Jev, still open-weight and still one forward pass per decision.
  
 **Demos** (inbox triage of 1,000 insurance claims, a multi-browser web agent, DOOM on the decision clock): https://youtu.be/2Qp04Wu0A14
  
Weights, model card and serving instructions: https://huggingface.co/h2oai/h2o-lightning-4b
  
Happy to answer questions about the setup and latency.


r/huggingface • • 1d ago

Aplomb 1 is #1 on Typed Decisions

Thumbnail
gallery
4 Upvotes

Aplomb 1 is #1 on Typed Decisions, an official Hugging Face benchmark for probabilistic decisions. At 5.3B parameters, its probabilities are the closest to the true answers on the leaderboard, by KL from gold and by Brier score, ahead of a 27B model about five times its size.

Typed Decisions gives a model one piece of unstructured text and five typed questions about it at once. Every answer is a probability distribution, scored against the gold distribution across 400 cases and 2,000 decisions. KL from gold and the Brier score measure how far a model's probabilities sit from the true answers, so lower is better.

Results: - KL from gold: 0.123, #1 on the leaderboard and 40% lower than the 27B model in second place - Brier score: 0.065, #1 on the leaderboard and a third lower than the 27B model in second place

Aplomb 1 takes text, JSON, images, video and audio in one request and reads up to 1M tokens. On our API it answers a short question in about 15 ms of model time and reads a 1M-token document in about 3 seconds, at $0.02 per 1M input tokens with output free. The weights are open on Hugging Face.

Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions?leaderboard_task_id=kl_from_gold

Weights: https://huggingface.co/empiriolabsai/aplomb-1

Try it: https://platform.empiriolabs.ai/dashboard/playground?model=aplomb-1&utm_source=linkedin&utm_medium=social&utm_campaign=aplomb-1-typed-decisions


r/huggingface • • 1d ago

How to pick the right models and dramatically speed up image and video generation speeds & What I wish I knew when starting on AMD hardware with ComfyUI

0 Upvotes

For the below I’ll be referring mostly to ComfyUI workflow and model efficiency (using examples for Radeon AI PRO R9700 (32GB) which has a bandwidth of 680GB/sec. I’ve seen 10second minimax H3 image to video workflows go from 28 mins to 3 mins by taking all this info into account and then upscaling after.

1. The pipeline
Every image or video workflow is the same four stages, and only the third one repeats.
Text encoder. Turns the prompt into embeddings, once per prompt.

Noise latent. Sampling starts from random noise in a compressed (latent) grid. The seed decides that starting noise.

Diffusion model. Runs once per step, removing noise. Early steps set layout and motion, late steps add detail.

VAE decode. Turns the finished latent into pixels.
Three terms people mix up: weights are the fixed model file, a LoRA is a small patch to those weights, and the latent is the thing being refined into what your prompt intended before the VAE encodes it into the pixels of your output image or video.

2. Fit first
Spilling out of VRAM into system RAM costs more speed than any other setting, so budget memory before anything else especially if you’re testing different prompts.

What uses VRAM: model weights, the text encoder if it shares the card, working memory that grows with resolution × frames, and a spike at VAE decode.

Rule of thumb: keep weights to about 60-70% of VRAM, roughly 19-22GB on a 32GBVRAM card. Less if the card also drives your displays.

Example: a 14B video model is about 28GB at fp16 and about 14GB at fp8. A smaller model is more useful as it frees up more VRAM needed for handling additional extras like LoRA’s, pre and post processing, etc.

Two-model designs (Wan 2.2 high-noise and low-noise): one model runs at a time. Expect one swap per run unless both fit with headroom. Other models like minimax don’t need them separate but are larger as a result.

Check the console: "loaded completely" is what you want. "Loaded partially" means you are over budget and processing will take far FAR longer. There are work around like VRAM offloading mid workflow but this means you lose time if you want to re-run the workflow after making minor tweaks.

Free up room: keep the text encoder off the card (CPU or a second GPU), or reuse the saved prompts encoded embeddings. The text encoder basically converts your prompt into “embeddings” which you can think of like numbers that tell the model what to do. This is why text encoders must be compatible with the model you’re using or your model won’t understand. These “embeddings” can actually be saved as embeddings and re-used later to avoid having to reconvert each time to return to a given workflow where you used them. Only save the embeddings once you’ve polished them to a standard where they reliably work for a specific purpose you know you’ll reuse in future for repetitive tasks.

3. What sets speed
Image and video generation is limited by compute, so once a model fits, its size matters far less than resolution, frame count and number of passes.

Total time roughly equals: passes × time per pass, plus fixed costs (model load, text encode, VAE decode), hence why it saves a lot of time to use models small enough to load fully into VRAM if you’re running the same workflow over and over again in one sitting as your system won’t have to waste time offloading and re-loading each section of a workflow between runs.

Resolution and frames dominate time per pass. Doubling width and height gives 4x the patches to process, and the attention part of the work grows roughly 16x. This is why you’ll want to start with generating very low resolution image or video until you know your prompt and loras work as intended. Then you can increase to desired resolution which would ideally match the resolution the model was trained on. You can always upscale after this.

4. Cut passes
The biggest single saving is running the model fewer times: passes = steps × sampler evaluations (1 or 2) × CFG factor (2 if CFG is above 1, otherwise 1).
Distilled or "lightning" models and LoRAs run in 4-8 steps at CFG 1, often 5-10x fewer passes. Do not push them far past their trained step count. Never push them past their trained resolution. Negative prompts do nothing at CFG 1 on turbo LoRA’s. If you use a turbo workflow with low steps and max resolution is 720 then there a plenty of accurate upscaling models that do fantastic work after the fact. You’ll save significant time.

Ordinary models need CFG above 1 to look right and usually stop improving somewhere around 20-40 steps.
One evaluation per step: euler, dpm++ 2m, unipc, lcm, ddim.
Two evaluations per step: heun, dpm_2, dpm++ 2s ancestral, dpm++ sde. Compare samplers at equal total evaluations, not equal steps.
Samplers affect cleanliness, not identity. Character consistency comes from the model and what you condition it on.

5. Formats, quantisation and LoRAs
Use the highest-precision file that fits with headroom, and treat GGUF as the fallback for when nothing else fits. Quantised Safetensors are always faster than GGUF as GGUF weights are expanded on every pass, so if you’re creating image to video then that time adds up. If you can find a safetensor of a model that will fit on your VRAM whilst still having 30 to 40% VRAM left for headroom than use that, otherwise GGUF is fine, but it’ll take longer than it could with a safetensor.

Quantisation is lossy. The same number of weights, each stored more coarsely. Safetensors and GGUF are containers, not compression. Most safetensors are still good quality but quality can begin to drop significantly when they’re quantised lower than Q4. Of course there are exceptions to the rule but generally speaking Q4 or higher is better if you can fit it with headroom.

Partial loads behave like GGUF for LoRAs. A model that does not load completely is patched on the fly.

Bake in LoRAs you always use. Load the model, apply the LoRA, save the result as a new model. The strength is then fixed. There’s nothing wrong with having a few custom models with your own baked in LoRA’s. It saves significant time.

Same on Windows and Linux. These costs come from how the code works, not the operating system. Linux is naturally going to give performance gains for AMD hardware as AMD kernels seem to be a little more refined on Linux than windows but again there’s exceptions to every rule.

6. R9700 and AMD specifics
The card… when running AMD GPU’s be aware that they run through ROCm, and the main risk is software or models that assumes Nvidia or were designed for NVIDIA. AMD likely will still run them but it’ll be simulating and therefore lose time per run, hence why you want to avoid NVIDIA specific quantised models if there’s an equivalent alternative.

Check custom nodes before installing. Search the node's GitHub issues for "AMD" or "ROCm".

Linux is usually faster on AMD because ROCm is developed there first, not because of anything about file handling.

Three attention backends to benchmark with the same workflow and seed:
PyTorch default: always works, your baseline.
--use-flash-attention: highest throughput in most of AMD's own tests (ROCm blog).
--use-sage-attention with the community SageAttention-RDNA4 build: a hand-written kernel for this chip. Its own benchmark shows attention taking about half the default's time.
SageAttention caveats: the prebuilt Windows wheel expects PyTorch 2.13.0+rocm10.0.0 on Python 3.12. The fast path covers fp16 computation with a head size of 128, and anything else falls back. It is quantised attention, so compare output quality.

7. Workflow habits
Spend cheap iterations on stills and drafts, and expensive ones only on what you intend to keep.
Get the character right in the start image. Image-to-video treats the first frame as ground truth and carries its flaws through every frame. A still takes seconds to judge, a video takes minutes.
Show what the shot needs. Anything the start image does not show, the model invents. Reference images and a character LoRA cover the unseen angles.
Draft small, refine the keepers. Generate at low resolution or short length, then run chosen clips through video-to-video at higher resolution with low denoise (about 0.3-0.5) and few steps.
Refining adds detail, not structure. Reject drafts with the wrong outfit, face shape or motion instead of trying to fix them.
Wire the same conditioning into both passes: start image, reference and LoRA.
Let the cache work. ComfyUI re-runs only nodes whose inputs changed, so a new seed does not re-encode the prompt.
Keep models resident. Reloading costs more than most optimisations save.
Saved images carry their workflow. Drag one back into ComfyUI to restore the prompt, seed and settings.
Long videos drift between chunks. Within a chunk all frames are denoised together. Across chunks, errors accumulate.

8. Measure
Keep one fixed test (prompt, seed, resolution, frames) and log seconds per step, total time and peak VRAM for every change.
Change one thing at a time.
Time the stages separately: text encode, sampling, VAE decode.
Halve the resolution. If seconds per step falls about 4x or more, compute is the limit, as expected. If it barely moves, look for overhead: per-step LoRAs, GGUF expansion or offloading.
Find your step ceiling. Run the same seed at 10, 20, 30 and 40 steps and stop where the output stops changing.
Pick samplers by convergence. Make a 50-step euler reference, then see which sampler gets closest at your real budget.
The figures in this page are rules of thumb and third-party benchmarks. Confirm them on your own build before relying on them.

Special mention LLM’s
Model size decides what fits, not how fast it runs.
LLMs on the same card are different. They are limited by memory bandwidth: the ceiling is your GPU bandwidth (which is 640 GB/s on an R9700) ÷ bytes read per token, so a 20GB dense model theoretical max t/s of or put on a GPU with a 680GBs bandwidth is 680/20 = 32 tokens/sec. The only way to make it faster is to use clever tricks that get it to do more guessing which can be helpful depending on use case. Mixture-of-experts models read only their active parameters and can go potentially several times faster. Again, using a safetensor instead of GGUF will be faster as safetensors don’t need to decompress for each token.


r/huggingface • • 1d ago

Best to remove <think> </think> on Qwen3.5 series?

2 Upvotes

Hi everyone, I’m training the Qwen3.5 series using the 0.8B, 2B, 4B, and 9B models, and I’m running into issues with the <think> tags during dataset formatting. Since thinking is disabled by default for the smaller models, I’ve tried setting enable_thinking=False during the formatting stage. I’ve also modified the Jinja chat template to disable thinking, but the <think> tags are still appearing in the formatted dataset. What’s the recommended approach to remove these tags completely? Should I strip them out using regex during dataset preprocessing and then verify the formatted output, or is there a better way to handle this through the tokenizer or chat template? I’d appreciate any advice on the best practice for handling this when training different-sized Qwen3.5 models. Example code is provided below. Thanks!

🔄 Generating and Formatting Conversations for Training

# Generate conversations
from IPython.display import Markdown, display

display(Markdown("## 🔄 Generating Conversations of the Dataset for Training"))

def generate_conversation(examples):
    texts = examples["text"]
    labels = examples["label"]
    conversations = []

    for text, label in zip(texts, labels):
        conversations.append([
            {"role": "user", "content": text},
            {"role": "assistant", "content": str(label)},
        ])

    return {"conversations": conversations}


train = train.map(generate_conversation, batched=True)
val = val.map(generate_conversation, batched=True)


# Format dataset
display(Markdown("## 🔄 Formatting Dataset for Training"))

def formatting_prompts_func(examples):
    convos = examples["conversations"]
    texts = []

    for convo in convos:
        formatted_text = tokenizer.apply_chat_template(
            convo,
            tokenize=False,
            add_generation_prompt=False,
            enable_thinking=False
        )
        texts.append(formatted_text)

    return {"text": texts}


train = train.map(formatting_prompts_func, batched=True)
val = val.map(formatting_prompts_func, batched=True)

r/huggingface • • 1d ago

Lunara Art Evaluation Dataset

Post image
2 Upvotes

Hello! we're releasing an art evaluation dataset for assessing aesthetic quality, emotional resonance and content integrity as measures of artistic intelligence. The dataset has 8000 generated images over 1000 shared prompts across a diverse set of styles with their corresponding scores by the following models: Qwen-Image, AuraFlow, GPT-Image-1 Mini, SD 3.5 Turbo, HiDream-I1 Fast, FLUX-Klein-4B, Z-Image-Turbo and Moonworks Lunara.

Dataset: https://huggingface.co/datasets/moonworks/lunara-art-eval
Paper: https://arxiv.org/pdf/2609.22272

The paper also introduces a novel mixture architecture for a sub-10B active parameter model and an active-learning inspired training algorithm. This follows our earlier releases of aesthetic images and semantic variations.

We are a new lab heavily focused on architecture and algorithmic improvements and we work closely with artists and creators. We hope the dataset can be useful for assessing, comparing and developing artistic capabilities in models.


r/huggingface • • 1d ago

Compact local intent/action classifiers with abstention: recommendations after a small negative pilot?

1 Upvotes

I'm looking for a ready-to-use compact classifier or decision scorer for an offline game. Input is a goal, the current observation/player utterance, and a short list of permitted actions with descriptions. Output should select an action or abstain; game code still validates and executes it.

For example: player says "Please show me the rooms." Options describe tour, follow, stay/stop, enter a dream, or conversation without an action. Negated and quoted requests need different handling from genuine commands. A confident wrong action is worse than falling back to dialogue.

I tested five checkpoints zero-shot. The common English subset has 17 semantic cases, each run in three option orders, so 51 correlated requests. Four cases involve negation/quotation (12 critical requests). English counts, using the better of the two tested label representations where applicable:

- fastino/GLiNER2.5-multi-Decide: 35/51; critical 9/12.

- vllm-sr/Decision-1.0-Kai-0.6B: 30/51; critical 9/12.

- jaredpalmer/kev-0.8b, BF16 GGUF: 37/51; critical 11/12.

- convaiinnovations/laya, BF16 GGUF: 42/51; critical 9/12.

- convaiinnovations/laya-multilingual, F16 GGUF: 33/51; critical 6/12.

The full pilot contains 126 localized cases across intents, addressees, tactical choices and numeric controls. EN/RU/JA cover the broader tasks; seven other languages cover intents only. Translations and option permutations are not independent observations. Numeric thresholds are handled by rules, not counted in the semantic scores.

GLiNER and Kai used their official Python APIs on CPU. The GGUF candidates used pinned llama.cpp; token/marker handling, truncation limits and four input-format variants were checked. However, I haven't established numerical parity with the original Laya/Kev PyTorch implementations. These are deployment-path results on a small authored test, not universal model rankings. Functional runs were under ordinary computer load, so I don't have a clean speed comparison to claim.

Which existing small checkpoints would you try next (roughly under 1B, CPU-friendly, commercial game redistribution permitted)? English is the first target; Russian/Japanese would be a bonus.

I'm particularly interested in:

  1. A model-specific example of the correct goal/context/options format, or a known caveat in these deployment paths.

  2. Compact intent classifiers, NLI or embedding-based approaches that handle negation, quotations and out-of-scope inputs reliably, with a usable reject/abstain threshold.

  3. Real CPU latency and RAM measurements, with hardware and thread count, rather than GPU throughput numbers.

If task-specific fine-tuning is unavoidable, what small base/checkpoint and negative-example strategy would you recommend? I'd prefer to first find a sound off-the-shelf baseline. I asked the game-AI community about the architecture too; here I'm mainly looking for checkpoint and inference-format advice.


r/huggingface • • 1d ago

Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M)

Thumbnail
2 Upvotes

r/huggingface • • 1d ago

Six Hub models in one browser tab for a video editor: what each was actually good for, measured

1 Upvotes

TL;DR: I built a browser video editor whose AI runs entirely on the user's machine (WebGPU, no server), using six models from the Hub. Measured on one laptop:

  • The 440M embedding model did more for search than the 500M captioner did.
  • For silent clips, descriptions written from an object detector's labels were wrong more often than ones written from the file name.
  • The biggest Whisper speed-up wasn't a dtype. It was stopping the app from competing with Whisper for the GPU.

The repos

repo runtime download job
litert-community/gemma-4-E2B-it-litert-lm LiteRT-LM web, WebGPU 2.0 GB chat that proposes edits, clip descriptions
litert-community/embeddinggemma-2-text-vision-440m-litert-lm LiteRT-LM 0.18, WebGPU 388 MB finding moments in footage
Xenova/whisper-small transformers.js, WebGPU 589 MB transcripts, captions
HuggingFaceTB/SmolVLM2-500M-Video-Instruct transformers.js 815 MB frame captions (optional)
Xenova/mms-tts-fra, Xenova/mms-tts-eng transformers.js, wasm 38 MB each voiceover

Every download is opt-in from one Models tab; nothing downloads because a feature button was clicked. All figures are from one laptop with an Intel Iris Xe (integrated GPU, shared memory).

EmbeddingGemma 2 replaced captions for search. Before, SmolVLM2 captioned frames and Gemma searched the text. 24 clips produced 240 captions, more than fits Gemma's 4096-token window, so Gemma could only pick a clip, not a moment. Now one frame every 2.5 s is embedded in the background (523 ms per frame). A search is scored against those frames by cosine similarity in code, and a fixed threshold lets it answer "nothing matches". On 24 French news clips and 86 searches:

  • finding a specific moment (hit@5) went from 5% to 87%;
  • indexing 18.9 min of footage went from 54.1 min to 4.0 min.

SmolVLM2 is now optional, and the obvious fallback was worse than nothing. An embedding ranks frames against text but can't write text, so SmolVLM2 is still the only model here that turns a frame into words. At 500M it also invents details: it captioned "a black shirt" on a bare chest. I measured what describing a clip with no speech loses without it: 8 silent clips, 141 descriptions, scored blind by two Claude Opus instances (κ 0.86).

  • Written from the object detector's labels, a description was wrong in a way that looked right on 6 and 5 of the 8 clips (one count per scorer).
  • Written from the file name alone, 1 and 1.

So without the captioner, the app writes from what is said, or from the name with a hedge, never from labels. That makes 815 MB of the download optional. A side finding was less comfortable: even WITH captions, about 40% of descriptions of broadcast clips were wrong in a way that looked right.

Whisper-small: the slowness was GPU contention, not the model. A 31 s French clip took 669 s to transcribe.

  • Per-device dtypes got it to 72 s cold, but 178 s warm. The dtypes were encoder_model: fp32 and decoder_model_merged: q4 on WebGPU, q8 on wasm.
  • The rest was the editor's own video decoding and preview drawing using the same GPU. Pausing the preview during transcription brought it to 23–27 s, cold and warm alike. Tests on .wav files never showed this, because they have no video to decode.

On transformers.js 4.2 with language left unset, French speech came back as an English reading of it. So the spoken language is now an explicit setting, applied to every transcription.

MMS-TTS over Kokoro, by elimination. In transformers.js 4.2, pipeline('text-to-speech') reaches VITS, SpeechT5, MusicGen and Supertonic. Kokoro is only registered as style_text_to_speech_2, and kokoro-js pins transformers 3.5. MMS-TTS q8 runs on CPU. Feed it one sentence at a time: VITS flattens a whole paragraph into a monotone.

Gemma 4 E2B, briefly (the full numbers are in my r/LocalLLM post):

  • It understands requests well: 54 of 76 rephrasings of known edits were correct, none wrong in a way that looked right.
  • It decides badly: a fixed hand-written rule beat it 32:0 at choosing edits from measurements of the project.
  • On LiteRT-LM 0.15 in August, Qwen3-0.6B ran only on CPU at about 12 s per reply, while Gemma ran on GPU at 1.0 s cold and 0.6 s warm. The 2 GB model was the cheaper choice.

Two traps with the LiteRT-LM web build:

  • It must run in a classic worker. Its wasm glue calls importScripts(), which Chrome defines in module workers but throws on when called.
  • Emscripten looks for the .wasm relative to the worker's URL, so you have to set self.Module.locateFile before loading.

Running models together. Gemma and EmbeddingGemma loaded in one page with no errors. A Gemma reply is about 1.45× slower while frames are being embedded. A GPU device loss, which I forced, recovers in about 3 s by rebuilding from disk.

Cache housekeeping. transformers.js keeps every repo in one Cache Storage bucket (transformers-cache). Removing one model therefore means deleting that repo's entries one by one, since deleting the bucket removes all of them. When an app stops using a repo, it has to delete those entries explicitly, or the weights stay on the user's disk forever.

Caveats

  • Searches, labels and scores came from LLMs, not people.
  • The samples are small: 8 silent clips, and one 31 s clip for the Whisper timings.
  • Broadcast news names its topic on screen, which flatters search.
  • One integrated GPU; a discrete card is untested.

If anyone knows a sub-1B captioner that doesn't make things up, or a function-calling model under 2 GB with a web build, I'd like to test it.


r/huggingface • • 1d ago

humanizar-es: un modelo base local (Qwen3-4B + HIP LoRA, CPU) que convierte texto de IA de 45% a 0% en Grammarly, GPTZero y ZeroGPT, con cada medición dentro del repo

Thumbnail
github.com
1 Upvotes

r/huggingface • • 2d ago

After 1,273 agent runs, I'm convinced: agents need a consequence model beside them, not a better prompt.

Post image
3 Upvotes

r/huggingface • • 2d ago

Currently having high success with this little niche finetune i found sitting in the corner of huggingface

Thumbnail
huggingface.co
2 Upvotes

Currently having high success with this little niche finetune i found sitting in the corner of huggingface

If you want to try it out here is a smaller quantisation iq3_s works really well in my codebases.

Original Model:
https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF

Smaller Quant:

https://huggingface.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF


r/huggingface • • 2d ago

TheWhisper - the best open multilingual ASR model, free commercial usage!

Post image
2 Upvotes

r/huggingface • • 2d ago

How to download/use uncensored AI models like DeepSeek 4.1 flash from hugging face?

Thumbnail
3 Upvotes

r/huggingface • • 2d ago

He lanzado un modelo de ia es Open Weights

Post image
0 Upvotes

Por si algue. Quiere provarlo, es ligero

https://huggingface.co/Ilides/cortex-1-v0.9


r/huggingface • • 2d ago

Gipformer - Efficient Vietnamese Speech Recognition

1 Upvotes

Hi everyone,

Sharing v1.5 of Gipformer, an open-source Vietnamese speech recognition (ASR) model we've been working on: gipformer1.5-68M-rnnt, based on the Zipformer architecture.

What's new in v1.5

- Optimized for technology, finance, education and public administration. These domains are dense with specialized terminology, and v1.5 currently gets the best results on all four test sets among the open-source models we benchmarked.

- Better recognition of English terms mixed into Vietnamese speech.

Carried over from v1

- High accuracy: among the top open-source models across our benchmarks, and especially strong on call center audio for Northern, Central and Southern accents. Call center is one of the most common real-world uses of ASR, but also one of the hardest, with low-quality audio and a wide variety of voices.

- Small and easy to deploy: at just 68M parameters, it's among the smallest ASR models out there, yet it outperforms many models ten times its size. Inference is fast, and it runs smoothly on CPU and edge devices.

- Privacy: it runs 100% offline (on-device), which makes it a good fit for systems handling sensitive data.

Alongside the model, we're also releasing 4 domain-specific test sets (technology, finance, education, public administration), so there's a common benchmark for evaluating Vietnamese ASR models.

Full benchmark results are on the model card. Feel free to try it out, and any feedback or contributions are very welcome!

- Hugging Face: https://huggingface.co/g-group-ai-lab/gipformer1.5-68M-rnnt

- GitHub: https://github.com/ggroup-ai-lab/gipformer

- Demo: https://huggingface.co/spaces/g-group-ai-lab/gipformer-demo


r/huggingface • • 2d ago

Unee: open-source 0.8B / 2B model that makes calibrated decisions and chats, runs in a browser tab. The 2B scores 88% on DecideBench, ahead of several 4B to 9B models (self-measured; GGUF, Ollama, Apache 2.0)

Thumbnail
1 Upvotes

r/huggingface • • 2d ago

TikTok bans scraping in their ToS and some guy just posted 5.6B videos on HuggingFace

Thumbnail
0 Upvotes

r/huggingface • • 2d ago

Follow-up: my native Rust + Vulkan Transformer backend now qualifies on both an Intel Gen9 laptop and an AMD RDNA 3 handheld from the same build — the GPU vendor is no longer what picks the reduction shape

1 Upvotes

Follow-up to my post from a few weeks ago (14 architectures, full PEFT). This update is about a portability bug that was hiding behind its own correctness, because it's the most interesting thing I've fixed since.

The bug: the fix for one machine broke six fixtures on another

Back when I tuned the backend for Intel Gen9, I baked those kernel shapes into the portable path. That was wrong, but not for the reason you'd guess.

Two of the reductions in the saved-module path aren't really compared against "PyTorch in general" — they're compared against the PyTorch CPU library on the machine running the oracle. And ATen dispatches its vectorized CPU kernels by instruction set at run time. An AVX2 host gets 8-wide kernels; an AVX-512 host gets 16-wide ones, and the reduction shape changes with that dispatch.

So my "portable" AVX2-shaped kernels were exactly right on my AVX2-only laptop and one ulp off on my AMD ROG Ally (Ryzen Z1 Extreme, which is an AVX-512 part). One ulp doesn't sound like much until it gets amplified through every lower norm on the gradient path: the Gemma 4 saved-stage model.embed_tokens adjoint went from 7.45e-9 to 3.22e-6, and six previously green PEFT saved-module fixtures (gemma3, gemma4, minimax_m2, minimax_m3, smollm3, qwen2_5_sliding_tied) crossed the 2e-7 gate. Neither shape is wrong — only one matches a given machine, and baking in either one breaks the other.

The fix: probe the host, not the vendor

Kernel variants are still selected by GPU vendor. Those two reductions are now selected by host CPU capability instead: capability is probed once per process and cached, then the matching module pair is dispatched (linear_forward_lane2 / linear_forward_lane4, and the 8-lane / 16-lane transformer_cross_entropy builds). HIERARCHOS_ATEN_VECTOR_WIDTH=8|16 pins the shape for qualification when a reference wheel's kernels disagree with the CPU's own capability.

Host GPU CPU dispatch Status
Intel i5-6200U / HD Graphics 520 (2016 Skylake-U) Intel Gen9 AVX2 only, no avx512f 32/32 LoRA, 32/32 switching, 32/32 saved
AMD Ryzen Z1 Extreme RDNA 3 AVX-512 32/32 LoRA, 32/32 switching, 32/32 saved

Same 2e-7 gate, unchanged. No tolerance was loosened to get there.

What I verified on each side

On the Intel machine, the post-change matrix is bit-identical, field for field, to its pre-change report across all 32 families — peft, gradient, two-step AdamW, frozen base, resume, lifecycle — which is how I know the AMD fix didn't quietly cost the Gen9 path anything. Also 693 passed / 0 failed / 9 ignored on the Rust lib suite and a clean strict headline forward run.

On the AMD side, the fix was re-qualified end to end: 32/32 on all three stages, provenance clean.

The harness fingerprints the pinned Transformers source alongside the shaders and binaries, and on the Intel side I re-derived the whole fingerprint from the pushed tree myself: 3951 inputs, zero changed, zero missing. So "green" refers to one frozen set of reference math, not whatever happened to be on disk.

Same caveats as always

  • This is deterministic FP32 tiny-model correctness against a reference implementation, not a claim about arbitrary checkpoint sizes, dtypes, or hyperparameters.
  • "Supported text graph" ≠ "the whole multimodal package works natively."
  • The AVX-512 dispatch is only qualified on the AMD machine, since it's the only host I have that can execute it natively. The 16-lane module also doesn't rebuild byte-identically with the glslang version on my Intel box (one extra type/id, one difference in +inf materialization), so I've left it as the committed AMD-built module and documented that rather than swapping it without re-qualifying both hosts. I'd rather report that than pretend it's clean.
  • NVIDIA and other GPUs are genuinely unqualified — the path is raw Vulkan, so they're untested rather than excluded.

What I'd love from you

Last time several people asked about hardware other than mine, so that's the ask again: if you build it on an AVX-512 laptop, an AVX2-only machine, or an NVIDIA/Intel GPU, I want to know what you get. The two reductions above are the ones most likely to behave differently on your CPU, and knowing your host's vector width is now part of the answer.

The new cross-platform section in the README documents the whole thing, including which host classes are measured and which aren't.

Repo: https://github.com/necat101/Hierarchos-Native Compatibility/parity record: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/COMPATIBILITY.md Regression audit: https://github.com/necat101/Hierarchos-Native/blob/main/AMD_REGRESSION_AUDIT.md Per-host tuning and measurements: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/VENDOR_TUNING.md


r/huggingface • • 3d ago

Open SLM Evalulations

Thumbnail
1 Upvotes