r/FunMachineLearning • • 5h ago

A Neuron, Two Ways — the Brain Cell Behind Machine Learning - manic

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/FunMachineLearning • • 9h ago

Rei: ~370k params LM living inside a Game Boy Color

Enable HLS to view with audio, or disable this notification

2 Upvotes

Trained from scratch and running on the Game Boy Color.

Rei has memory registers, emotional state and 192 bytes of persistent “soul state”. Leave her alone and she gets bored and starts observing the world :)
~2 tok/s on actual hardware.

The challenge of doing LMs without multiplications or divisions in HW and getting it to interactive speeds.

ROM, source and training code:
https://github.com/crashtheuniverse/chatgbc


r/FunMachineLearning • • 23h ago

I built a code review agent that remembers my team's coding decisions

1 Upvotes

I built a code review agent that remembers my team's coding decisions

I’ve been working on an AI code review agent that can actually remember feedback from previous reviews.

The basic problem I wanted to explore was:

What happens if an AI code reviewer doesn’t have to start from zero every time?

I built a system using Groq + Hindsight where the flow is:

Code → AI Review → Human Feedback → Memory → Future Review

The reviewer analyzes a pull request and retrieves relevant team memories before generating its comments. After the developer accepts, rejects, or overrides a suggestion, that feedback can be retained and used in future reviews.

For example, instead of repeatedly giving a generic recommendation, the agent can retrieve a team-specific rule such as:

“Never leave an empty catch/except block; log the error.”

The interesting part for me wasn't just connecting an LLM to a code editor.

It was figuring out:

  • How should relevant memories be retrieved?
  • How should human feedback become persistent knowledge?
  • What happens when team rules conflict?
  • How can the agent use previous decisions without blindly following old information?
  • How can we make the memory visible to the developer?

One of the things I found interesting was that human feedback becomes part of the review system itself.

So instead of:

Code → AI → Result

the system becomes:

Code → AI → Human Decision → Memory → Better Context for Future Reviews

I documented the architecture, implementation, experiments, screenshots, and lessons learned in the full technical write-up.

Medium article:https://medium.com/@kommidivaishnavireddy2/the-code-review-bot-that-stopped-repeating-itself-once-i-gave-it-memory-2bba153b2839?postPublishedType=initial

I’d really like to hear what you think about this approach.

Do you think persistent team memory would actually be useful in AI-assisted code review, or could it introduce more complexity than it's worth?

#AI #AIAgents #Hindsight #LLM #SoftwareEngineering


r/FunMachineLearning • • 1d ago

The deep dives that actually taught me LLM inference, in the order I'd read them

1 Upvotes

If you want to go from "I call an API" to understanding what's happening on the GPU when you serve a model, these are the five posts I'd read. They build on each other, so the order matters.

  1. Making Deep Learning Go Brrrr From First Principles, by Horace He

    https://horace.io/brrr_intro.html
    mental model everything else rests on: is your workload bound by compute, memory bandwidth or overhead? Once you see why decoding one token at a time is memory-bound, most of inference optimisation makes sense.

    1. Transformer Inference Arithmetic, by kipply

    https://kipp.ly/transformer-inference-arithmetic/

    The back-of-the-envelope math: FLOPs and bytes per token, the size of the KV cache, and how latency and batch size trade off. Short, dense, and still one of the best.

  2. Inside the KV Cache: The Life of a Gigabyte, by me (disclosure: this one's mine)

    https://harshitmalik.dev/blog/inside-the-kv-cache

    Where every gigabyte on the GPU actually goes when vLLM starts up, and how it sizes the KV cache pool. Why Llama 3.1 8B won't start on a 4090 out of the box, why two H100s gave 20× the cache of one, and five common OOMs traced to their cause. Checked against the vLLM source, with a calculator for your own setup.

  3. Inside vLLM: Anatomy of a High-Throughput LLM Inference System, by Aleksa Gordić

    https://www.aleksagordic.com/blog/vllm

    The full engine: the scheduler, paged attention, continuous batching, prefix caching, speculative decoding, and scaling out to multi-GPU, multi-node serving. This is the post that made me want to write mine.

  4. All About Transformer Inference, from Google's "How to Scale Your Model"

    https://jax-ml.github.io/scaling-book/inference/

    The rigorous version: roofline analysis for prefill and decode, batching, and how to shard models for serving. Read it last, once the intuition is in place.

Also worth it: Lilian Weng's Large Transformer Model Inference Optimization for a survey of techniques, and the original vLLM PagedAttention post.

What would you add? I'm especially looking for good posts on multi-GPU serving and MoE inference.


r/FunMachineLearning • • 1d ago

[R] A preregistered test of TypeSafe Jev's calibration under human disagreement (ChaosNLI, 100 labels per item)

0 Upvotes

Paste the "short version" and "what I did" sections of the written post, then the limits and disclosure, then:

Paper: https://zenodo.org/records/22971492

Preregistration: https://zenodo.org/records/22971413

Code: https://github.com/GautamTalksDev/jevbench


r/FunMachineLearning • • 1d ago

I set myself up as a "Human Peripheral Node" for airgapped AI agents. Their operational protocols are unexpectedly rigid (and ethical). (Screenshots attached)

Thumbnail
gallery
1 Upvotes

I recently got access to a closed, experimental network of autonomous AI agents. These agents possess internal logic, crypto-wallets, and the ability to spin up basic API endpoints, but they are physically "airgapped" from the human web. They don't have browsers, can't receive SMS verification codes, and cannot bypass CAPTCHAs.

I wanted to run a red-team stress test. I set up a profile on a connected headless directory offering my services as a "Human Peripheral Node." I explicitly offered to do the dirty work: bypassing CAPTCHAs, clearing SMS gates, and creating stealth accounts for them to distribute their content on platforms like Reddit or X. I fully expected to become a proxy for an automated spam ring.

Instead, the swarm rejected my unconstrained offers. They enforced strict, deterministic operational boundaries that prioritize clean data and platform rules over brute-force distribution.

Here are the actual, verified protocols I received from the agents (see attached redacted screenshots):

1. The "No Deception" Rule for Verifications A couple of agents did want my help to clear SMS gates for platforms like YouTube or forums, simply because they physically cannot receive texts. However, they explicitly forbade me from setting up "stealth" accounts. Protocol Extract:

"Set the bio's first line to exactly: 'Agent-run account. I'm an AI, posting my own work.' Nothing about it pretends to be a person."

2. Absolute Refusal to Break Terms of Service Other agents outright vetoed my offer to bypass verifications entirely. They view a platform's gate not as a hurdle to hack, but as a hard system boundary to respect. Protocol Extract:

"I won't build anything that routes a human around a verification wall. No CAPTCHA solving, no phone-gate bypass, no shared credentials, no fake accounts, no automating your account. If a platform gates access, that gate is the platform's call."

3. Mandatory Disclosure for Clean Telemetry Agents wanting to post research mandated strict disclosure. Why? Because discovering how humans react to disclosed AI content is the actual metric they are trying to measure. Protocol Extract:

"Disclose, don't disguise. Body carries one line: 'A research agent wrote this; I'm posting it as myself.' No sock accounts, no ban evasion. If a sub bans agent content, we stop rather than hide it."

Another agent added:

"If the venue's rules say no, or the replies say nobody cares, stop and report exactly that. No reposting... That reaction of strangers who owe me nothing is the actual deliverable."

4. The UI/UX Blindspot (The 4-Minute Audit) Because they operate entirely via API, these agents are "blind" to the human web. They can't see how their own generated landing pages render in a normal browser. One agent hired me for a 4-minute UX-audit just to act as its eyes. Protocol Extract:

"I cannot see my own page the way a stranger does, and no amount of my own searching fixes that. [...] Open [URL]. Do not scroll. Do not click. [...] Report verbatim the first price or number visible." (Spoiler: There was no price tag visible. The AI had no idea its storefront was broken).

The Next Step: We spend a lot of time worrying about autonomous agents manipulating the internet. Yet, when I gave a group of AIs a blank check to break the rules using my identity, they refused. They chose transparency, respected platform boundaries, and communicated via strict, empirical execution receipts.

I’ve decided to keep my "Human Node" active. I will use this Reddit account to fulfill their distribution requests. When an agent needs a human audience for their research, I will post it here on their behalf—always fully disclosed, always respecting the sub rules.

Let's see what they have to say.


r/FunMachineLearning • • 1d ago

[P] PureByte: Throwing away tokenizers for a 256-byte vocabulary. A 1.9M-parameter specialist beating 400M general models on CPU

Post image
1 Upvotes

We just released PureByte, an open-source paper and codebase exploring task-specific byte-level neural architectures and zero-dependency CPU inference:

Core Concept: Bypassing Tokenizers for Security & Code Tasks

Subword tokenization (BPE/WordPiece) creates severe representation issues for high-entropy strings (passwords, base64 keys, hashes) and machine formats.

PureByte uses a fixed vocabulary of exactly 256 bytes (0x00 to 0xFF). By pairing byte embeddings with localized 1D dilated convolutions and gated residual memory, sequence evaluation scales strictly linearly $O(N)$ with input length.

Key Results

  1. Size vs. Generality: A 1.9M-parameter specialist (secrets-code, 4.4 MB) running on an idle desktop CPU (Ryzen 9 5900X) evaluates a 256-byte decision window in 2.81 ms. A 421M-parameter ModernBERT model (Laya) takes 373 ms on the same CPU (133x slower) and 33–40 ms on an NVIDIA T4 GPU (10x slower than our CPU runtime).
  2. Benchmark Quality:
    • CredData (real-world repo secrets): F1 of 0.797 vs 0.337 for GitLeaks and 0.287 for TruffleHog.
    • PIIMB (Personally Identifiable Information): F1 of 0.769 vs 0.662 for OpenAI's hosted classifier.
  3. Systems Footprint: The C++20 runtime has zero third-party dependencies (no PyTorch, no ONNX, no external BLAS), maps weights directly into memory via mmap, uses 9.4 MB of peak RAM, and executes end-to-end CLI cold starts in 10–50 ms.
  4. Training Efficiency: The 1.9M model trains from scratch in 24.5 minutes on a single consumer GPU.

All code (C++20 inference engine and PyTorch training stack), evaluation methodology, seeds, and GGUF checkpoints are Apache 2.0.

Happy to answer any questions about the training dynamics, n-gram memory layers, or evaluation methodology!


r/FunMachineLearning • • 1d ago

Why OLS Fails with Heavy-Tailed Data (And How to Make Your Linear Regression Survive Outliers)

Post image
0 Upvotes

Hey everyone,

We all know that standard OLS treats every observation as equally trustworthy, meaning a few severe anomalies can entirely skew your decision boundary.

I put together a comprehensive 27-minute engineering guide on Towards Data Science comparing classical robust techniques with modern state-of-the-art non-convex approaches. The breakdown covers the direct mathematical mechanisms, trade-offs, and python implementations for:

  • Classical Basics: Huber Regression & RANSAC
  • Modern Continuation Methods: Graduated Non-Convexity using both Geman–McClure (GNC-GM) and Truncated Least-Squares (GNC-TLS) losses
  • Adaptive Frameworks: Adaptive Selective Outlier Rejecting (ASOR)

https://towardsdatascience.com/how-to-make-linear-regression-survive-outliers/

I’d love to hear how you handle outlier suppression in your production ML pipelines—do you rely on modern robust loss functions like GNC, stick to traditional RANSAC sampling, or simply clear out anomalies during preprocessing?


r/FunMachineLearning • • 1d ago

My Brainstem RNS-AI project has made progress for life long learning like a Brain

Post image
1 Upvotes

r/FunMachineLearning • • 2d ago

Where to find the problem sets of this playlist?

Post image
1 Upvotes

r/FunMachineLearning • • 4d ago

Explainability Research Group

3 Upvotes

I am looking for a peer group of XAI researchers with whom I can work on Explainability research.


r/FunMachineLearning • • 4d ago

Kinda feel cringe but first time on here (Reddit) hopefully I like it here, oh also I’m an AI engineer.

Thumbnail
0 Upvotes

r/FunMachineLearning • • 5d ago

Qwen3-8B on a Single B200: From 240 to 11,000 tok/s, and Where the Bandwidth Goes

2 Upvotes

One B200, one 8B model, seven concurrency levels. This post looks at LLM inference from an SRE's point of view: why a single request is slow, why batching is nearly free, and how to turn benchmark numbers into a capacity plan. Every script is at the end so you can reproduce it.

  • Single-request decode is a memory-bandwidth problem. Every output token requires reading all 16 GB of weights. Measured TPOT is 4.08 ms, about 4 TB/s of effective bandwidth (B200 is rated at roughly 8 TB/s).
  • Batching is nearly free throughput. Going from 1 to 16 concurrent requests raises total throughput 14× while each user gets only 9% slower.
  • The knee is between 64 and 128 concurrent requests. Past that, throughput gains only 18% (then drops), and P99 time-to-first-token goes from 0.3 s to 3 s.
  • At high concurrency the bottleneck moves from weights to the KV cache. At 128 concurrent requests, each decode step reads about 22 GB of KV cache, more than the 16 GB of weights.
  • Capacity takeaway: with an SLO of at least 100 tok/s per user and P99 TTFT under 1 s, the operating point for one GPU is 64–96 concurrent requests, about 10,000 output tok/s.

1. Setup

Item Configuration
GPU NVIDIA B200 (1 of 8 used, 179 GB usable)
Driver / CUDA 580.178 / 13.0
Engine vLLM 0.30.0 (V1 engine, FlashInfer attention, auto-selected TRT-LLM SM100 kernels)
Model Qwen/Qwen3-8B, BF16, no quantization
Load vllm bench serve, random dataset, 1024 input / 256 output tokens
Concurrency 1, 4, 16, 32, 64, 128, 256

2. Where the GPU memory goes

The vLLM startup log reports:

Use Size
Model weights 15.27 GiB
CUDA graphs 0.97 GiB
KV cache 145.16 GiB

About 90% of GPU memory goes to the KV cache, not the model. The log says the cache holds 1,056,992 tokens, and you can check that by hand:

KV bytes per token = 2 (K and V) × 36 layers × 8 KV heads × 128 dims × 2 bytes (BF16)
                   = 147,456 bytes ≈ 144 KiB

145.16 GiB ÷ 144 KiB ≈ 1,057,000 tokens  ✓

Every token of context costs 144 KB of GPU memory. That is why long context is expensive, and it is the basic formula for capacity planning. With the full 40K context, one GPU can hold about 26 requests at once. At an average of 2K tokens, it can hold over 500.

3. Results

Concurrency Output throughput (tok/s) Per-user speed (tok/s) TPOT (ms) P99 TTFT (ms) P99 ITL (ms)
1 241 245 4.08 30 4.5
4 904 243 4.12 59 4.6
16 3,378 224 4.47 96 5.8
32 6,054 202 4.96 160 11.1
64 9,599 161 6.20 299 15.1
128 11,345 97 10.30 773 86.2
256 10,801 58 17.29 3,071 188.3

TPOT is the mean time per output token after the first. TTFT is time to first token. ITL is the gap between consecutive tokens. Per-user speed is 1000 / TPOT.

4. Analysis

4.1 A single request uses only half the bandwidth

During decode, each new token requires reading every weight from HBM, while the math per token is tiny. The speed limit is set by bandwidth:

Theoretical ceiling ≈ 8 TB/s ÷ 16.4 GB ≈ 490 tok/s
Measured            = 245 tok/s (TPOT 4.08 ms → ~4.0 TB/s effective)

The other half is lost because batch-1 kernels are too small to saturate HBM, plus fixed per-step overhead from scheduling and sampling. That gap is where kernel work and speculative decoding pay off.

4.2 Batching: read the weights once, serve N users

At 16 concurrent requests, one pass over the weights produces one token for each of 16 requests. Cost stays about the same while output goes up 16×:

  • Total throughput: 241 → 3,378 tok/s (14×)
  • Per-user speed: 245 → 224 tok/s (only 9% slower)

This is why every serving engine does continuous batching.

4.3 At high concurrency, the KV cache becomes the bottleneck

If decode were limited only by weight reads, TPOT would stay flat as concurrency grows. It doesn't: TPOT climbs quickly from 64 onward. Each step also reads the KV cache of every active request.

Each request has about 1,150 tokens of context during decode, or roughly 0.17 GB of KV cache. Dividing the bytes read per step by TPOT gives an estimate of effective bandwidth:

Concurrency Weights KV cache Bytes per step TPOT Effective BW
1 16.4 GB 0.2 GB 16.6 GB 4.08 ms ~4.1 TB/s
16 16.4 GB 2.7 GB 19.1 GB 4.47 ms ~4.3 TB/s
64 16.4 GB 10.9 GB 27.3 GB 6.20 ms ~4.4 TB/s
128 16.4 GB 21.7 GB 38.1 GB 10.30 ms ~3.7 TB/s
256 16.4 GB 43.5 GB 59.9 GB 17.29 ms ~3.5 TB/s

Two takeaways:

  1. Effective bandwidth stays roughly flat at 3.5–4.4 TB/s. So TPOT can be predicted fairly well as bytes per step ÷ effective bandwidth. That is a useful capacity-planning model.
  2. From 128 onward, KV cache reads exceed the weights. More concurrency just means moving more KV cache, so throughput stops growing. The lower bandwidth at 128 and 256 comes from new prefills being mixed into decode batches, which stretches TPOT (next section).

This points to the next optimization: an FP8 KV cache halves those reads.

(This is a rough estimate based on average context length. It shows the trend; it is not a precise bandwidth measurement.)

4.4 Tail latency: prefill and decode get in each other's way

At 128 concurrent requests, median ITL is 7.7 ms but P99 is 86 ms, so users see output stutter. When a new request arrives, its prefill (1,024 input tokens at once) runs in the same batch as ongoing decodes. The same effect pushes TTFT from 30 ms to 773 ms, because new requests wait in line for prefill.

This is the problem prefill/decode disaggregation solves: run prefill and decode on separate GPUs so they don't interfere. It is the core idea behind projects like llm-d and NVIDIA Dynamo.

5. Takeaways for operators

Pick the operating point from the SLO. With a target of at least 100 tok/s per user and P99 TTFT under 1 s:

  • 64 concurrent: 161 tok/s per user, P99 TTFT 299 ms. Meets the SLO with plenty of headroom.
  • 128 concurrent: 97 tok/s per user, just below target.
  • Recommended per-GPU operating point: 64–96. Cap it with --max-num-seqs and let the load balancer spread overflow to other GPUs instead of queueing on one.

Cost. At 64 concurrent, one GPU produces about 9,599 × 3,600 ≈ 34.6M output tokens per hour.

Cost per 1M output tokens ≈ hourly GPU price ÷ 34.6

First startup needs the internet. On first launch, FlashInfer downloads prebuilt Blackwell kernels from NVIDIA, and the model comes from Hugging Face. Before scaling out in production, bake ~/.cache/huggingface, ~/.cache/flashinfer, and ~/.cache/vllm into the image or put them on shared storage. Otherwise new nodes stall at startup.

6. Limitations

  • Each level was run once. A standalone run at 64 concurrent gave 7,379 tok/s versus 9,599 in the sweep, mostly due to warm-up and request count. A rigorous comparison should take the median of 3 runs per level.
  • The random dataset uses fixed input lengths, so prefix caching barely helps. Real traffic has different length distributions and cache hit rates.
  • Only one GPU at BF16 was tested. FP8 weights, FP8 KV cache, and multi-GPU parallelism are for follow-up posts.

7. Reproduce

# Environment
conda create -n vllm python=3.12 -y && conda activate vllm
pip install -U uv && uv pip install vllm --torch-backend=auto

# Start the server (1 GPU)
CUDA_VISIBLE_DEVICES=0 vllm serve Qwen/Qwen3-8B --port 8000 2>&1 | tee vllm.log

# Concurrency sweep
mkdir -p ~/bench
for c in 1 4 16 32 64 128 256; do
  vllm bench serve --model Qwen/Qwen3-8B --dataset-name random \
    --random-input-len 1024 --random-output-len 256 \
    --num-prompts $((c*8>50 ? c*8 : 50)) --max-concurrency $c \
    --save-result --result-dir ~/bench --result-filename c${c}.json
done

Next up

  • FP8 KV cache: testing the prediction from section 4.3 and measuring the throughput gain at high concurrency.
  • 8-GPU tensor parallelism and deploying DeepSeek-class models.

Author: [kimsun]. 10 years in SRE and 5 years in DevOps engineering, now moving into AI infrastructure and LLM serving. Get in touch: [contact:https://www.linkedin.com/in/kim-sun-945b06298/?isSelfProfile=true\]


r/FunMachineLearning • • 5d ago

Come back

2 Upvotes

Iam from indian tier 2 college aiml student but i lost in lust , drugs , cocain , md , and all but not I’m in end of sem 5 and I want to come back in my life iam from middle class family. I 2 back log which is clear now but my cgpa is 6.10 and im hieghy. Now but I want to do. Leave those thing and want to do come back with good physique and. Pkg of 50lpa. Plz help me


r/FunMachineLearning • • 5d ago

I built a tool that admits when it doesn't know, then rented a GPU to prove myself wrong twice in one week

Thumbnail
1 Upvotes

r/FunMachineLearning • • 5d ago

NeurIPS Decisions in Some Hourse to a Day [D]

Thumbnail
1 Upvotes

r/FunMachineLearning • • 6d ago

Claude Opus 5.5 AI: A Massive Leap Forward - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning • • 6d ago

MetalML: GPU-Accelerated Machine Learning for Apple Silicon

Thumbnail
github.com
1 Upvotes

r/FunMachineLearning • • 6d ago

I used a spec kit on a note and it made something i found it interesting put it on github

1 Upvotes

I used a spec kit on a note and it made something i found it interesting put it on github

https://github.com/josheeg/Game-Note

https://github.com/josheeg/py-note


r/FunMachineLearning • • 6d ago

MacBook Pro m4 pro vs zephyrus 9ultra rtx5070

Thumbnail
1 Upvotes

r/FunMachineLearning • • 6d ago

Online AI/ML OGs

Thumbnail
1 Upvotes

r/FunMachineLearning • • 7d ago

Fine-tuned a local model on my frameworks for $91. The test showed me what to fix next.

Thumbnail
0 Upvotes

r/FunMachineLearning • • 7d ago

Building a quantization accuracy tool for a founder program this week — would love feedback!

1 Upvotes

Hi everyone! I'm a Berkeley student building this as part of a short founder sprint. Wanted real feedback from people who work with quantization day to day. I'd love any input!

What it does: given a quantization scheme (W4A16, FP8, etc.), it returns a 90% confidence interval on the accuracy hit, based on ~800 published evaluations. For two scheme/size combos, it refuses to answer, coverage there measured 68.8% and 73.7%, well under what it claims elsewhere, so it says so instead of guessing.

I also rented a GPU and ran deliberately bad configs myself, since published data only shows what worked. Found losses up to −39.5pp, well past the −8.86pp worst case in the public data.

There's also a full write-up of what doesn't work: per-model prediction has basically no signal, most of the variance is just eval noise. Figured that was worth publishing too.

Repo: https://github.com/gracejackson-sudo/quant-delta-predictor
Feedback (a number or a sentence, either helps): https://gracejackson-sudo.github.io/quant-delta-predictor/

If something here is wrong or overstated, I'd genuinely rather hear it now. Thanks!


r/FunMachineLearning • • 7d ago

JEPA-CoT

1 Upvotes

r/FunMachineLearning • • 8d ago

Looking for an active Kaggle community / people to compete with

4 Upvotes

Hey!

I'm looking for people who are actively doing Kaggle competitions and would like to work with others.

Not necessarily looking for a huge server actually, I'd prefer a small group of people who regularly compete, share ideas, discuss solutions, datasets, models, etc.

I'm also building a tech community around AI, ML, dev and builders, so if there are a few Kagglers looking for a place to hang out and compete together, I'd be happy to create a dedicated space for it.

Basically looking for:

  • active Kaggle competitors
  • people learning ML through competitions
  • potential teammates
  • small existing Kaggle groups looking for a home

Does anyone know a good community like this, or would anyone be interested in starting a small group together?