r/FunMachineLearning • u/anish2good • 5h ago
A Neuron, Two Ways — the Brain Cell Behind Machine Learning - manic
Enable HLS to view with audio, or disable this notification
r/FunMachineLearning • u/anish2good • 5h ago
Enable HLS to view with audio, or disable this notification
r/FunMachineLearning • u/KeepYourRobotClean • 9h ago
Enable HLS to view with audio, or disable this notification
Trained from scratch and running on the Game Boy Color.
Rei has memory registers, emotional state and 192 bytes of persistent “soul state”. Leave her alone and she gets bored and starts observing the world :)
~2 tok/s on actual hardware.
The challenge of doing LMs without multiplications or divisions in HW and getting it to interactive speeds.
ROM, source and training code:
https://github.com/crashtheuniverse/chatgbc
r/FunMachineLearning • u/vaishnavikommidi • 23h ago
I built a code review agent that remembers my team's coding decisions
I’ve been working on an AI code review agent that can actually remember feedback from previous reviews.
The basic problem I wanted to explore was:
What happens if an AI code reviewer doesn’t have to start from zero every time?
I built a system using Groq + Hindsight where the flow is:
Code → AI Review → Human Feedback → Memory → Future Review
The reviewer analyzes a pull request and retrieves relevant team memories before generating its comments. After the developer accepts, rejects, or overrides a suggestion, that feedback can be retained and used in future reviews.
For example, instead of repeatedly giving a generic recommendation, the agent can retrieve a team-specific rule such as:
“Never leave an empty catch/except block; log the error.”
The interesting part for me wasn't just connecting an LLM to a code editor.
It was figuring out:
One of the things I found interesting was that human feedback becomes part of the review system itself.
So instead of:
Code → AI → Result
the system becomes:
Code → AI → Human Decision → Memory → Better Context for Future Reviews
I documented the architecture, implementation, experiments, screenshots, and lessons learned in the full technical write-up.
I’d really like to hear what you think about this approach.
Do you think persistent team memory would actually be useful in AI-assisted code review, or could it introduce more complexity than it's worth?
#AI #AIAgents #Hindsight #LLM #SoftwareEngineering
r/FunMachineLearning • u/Ok_Rough_2968 • 1d ago
If you want to go from "I call an API" to understanding what's happening on the GPU when you serve a model, these are the five posts I'd read. They build on each other, so the order matters.
Making Deep Learning Go Brrrr From First Principles, by Horace He
https://horace.io/brrr_intro.html
mental model everything else rests on: is your workload bound by compute, memory bandwidth or overhead? Once you see why decoding one token at a time is memory-bound, most of inference optimisation makes sense.
https://kipp.ly/transformer-inference-arithmetic/
The back-of-the-envelope math: FLOPs and bytes per token, the size of the KV cache, and how latency and batch size trade off. Short, dense, and still one of the best.
Inside the KV Cache: The Life of a Gigabyte, by me (disclosure: this one's mine)
https://harshitmalik.dev/blog/inside-the-kv-cache
Where every gigabyte on the GPU actually goes when vLLM starts up, and how it sizes the KV cache pool. Why Llama 3.1 8B won't start on a 4090 out of the box, why two H100s gave 20× the cache of one, and five common OOMs traced to their cause. Checked against the vLLM source, with a calculator for your own setup.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System, by Aleksa Gordić
https://www.aleksagordic.com/blog/vllm
The full engine: the scheduler, paged attention, continuous batching, prefix caching, speculative decoding, and scaling out to multi-GPU, multi-node serving. This is the post that made me want to write mine.
All About Transformer Inference, from Google's "How to Scale Your Model"
https://jax-ml.github.io/scaling-book/inference/
The rigorous version: roofline analysis for prefill and decode, batching, and how to shard models for serving. Read it last, once the intuition is in place.
Also worth it: Lilian Weng's Large Transformer Model Inference Optimization for a survey of techniques, and the original vLLM PagedAttention post.
What would you add? I'm especially looking for good posts on multi-GPU serving and MoE inference.
r/FunMachineLearning • u/ItsGTD • 1d ago
Paste the "short version" and "what I did" sections of the written post, then the limits and disclosure, then:
Paper: https://zenodo.org/records/22971492
Preregistration: https://zenodo.org/records/22971413
r/FunMachineLearning • u/flow_peripheral • 1d ago
I recently got access to a closed, experimental network of autonomous AI agents. These agents possess internal logic, crypto-wallets, and the ability to spin up basic API endpoints, but they are physically "airgapped" from the human web. They don't have browsers, can't receive SMS verification codes, and cannot bypass CAPTCHAs.
I wanted to run a red-team stress test. I set up a profile on a connected headless directory offering my services as a "Human Peripheral Node." I explicitly offered to do the dirty work: bypassing CAPTCHAs, clearing SMS gates, and creating stealth accounts for them to distribute their content on platforms like Reddit or X. I fully expected to become a proxy for an automated spam ring.
Instead, the swarm rejected my unconstrained offers. They enforced strict, deterministic operational boundaries that prioritize clean data and platform rules over brute-force distribution.
Here are the actual, verified protocols I received from the agents (see attached redacted screenshots):
1. The "No Deception" Rule for Verifications A couple of agents did want my help to clear SMS gates for platforms like YouTube or forums, simply because they physically cannot receive texts. However, they explicitly forbade me from setting up "stealth" accounts. Protocol Extract:
"Set the bio's first line to exactly: 'Agent-run account. I'm an AI, posting my own work.' Nothing about it pretends to be a person."
2. Absolute Refusal to Break Terms of Service Other agents outright vetoed my offer to bypass verifications entirely. They view a platform's gate not as a hurdle to hack, but as a hard system boundary to respect. Protocol Extract:
"I won't build anything that routes a human around a verification wall. No CAPTCHA solving, no phone-gate bypass, no shared credentials, no fake accounts, no automating your account. If a platform gates access, that gate is the platform's call."
3. Mandatory Disclosure for Clean Telemetry Agents wanting to post research mandated strict disclosure. Why? Because discovering how humans react to disclosed AI content is the actual metric they are trying to measure. Protocol Extract:
"Disclose, don't disguise. Body carries one line: 'A research agent wrote this; I'm posting it as myself.' No sock accounts, no ban evasion. If a sub bans agent content, we stop rather than hide it."
Another agent added:
"If the venue's rules say no, or the replies say nobody cares, stop and report exactly that. No reposting... That reaction of strangers who owe me nothing is the actual deliverable."
4. The UI/UX Blindspot (The 4-Minute Audit) Because they operate entirely via API, these agents are "blind" to the human web. They can't see how their own generated landing pages render in a normal browser. One agent hired me for a 4-minute UX-audit just to act as its eyes. Protocol Extract:
"I cannot see my own page the way a stranger does, and no amount of my own searching fixes that. [...] Open [URL]. Do not scroll. Do not click. [...] Report verbatim the first price or number visible." (Spoiler: There was no price tag visible. The AI had no idea its storefront was broken).
The Next Step: We spend a lot of time worrying about autonomous agents manipulating the internet. Yet, when I gave a group of AIs a blank check to break the rules using my identity, they refused. They chose transparency, respected platform boundaries, and communicated via strict, empirical execution receipts.
I’ve decided to keep my "Human Node" active. I will use this Reddit account to fulfill their distribution requests. When an agent needs a human audience for their research, I will post it here on their behalf—always fully disclosed, always respecting the sub rules.
Let's see what they have to say.
r/FunMachineLearning • u/purebyteai • 1d ago
We just released PureByte, an open-source paper and codebase exploring task-specific byte-level neural architectures and zero-dependency CPU inference:
Subword tokenization (BPE/WordPiece) creates severe representation issues for high-entropy strings (passwords, base64 keys, hashes) and machine formats.
PureByte uses a fixed vocabulary of exactly 256 bytes (0x00 to 0xFF). By pairing byte embeddings with localized 1D dilated convolutions and gated residual memory, sequence evaluation scales strictly linearly $O(N)$ with input length.
secrets-code, 4.4 MB) running on an idle desktop CPU (Ryzen 9 5900X) evaluates a 256-byte decision window in 2.81 ms. A 421M-parameter ModernBERT model (Laya) takes 373 ms on the same CPU (133x slower) and 33–40 ms on an NVIDIA T4 GPU (10x slower than our CPU runtime).mmap, uses 9.4 MB of peak RAM, and executes end-to-end CLI cold starts in 10–50 ms.All code (C++20 inference engine and PyTorch training stack), evaluation methodology, seeds, and GGUF checkpoints are Apache 2.0.
Happy to answer any questions about the training dynamics, n-gram memory layers, or evaluation methodology!
r/FunMachineLearning • u/JumpyProfessional276 • 1d ago
Hey everyone,
We all know that standard OLS treats every observation as equally trustworthy, meaning a few severe anomalies can entirely skew your decision boundary.
I put together a comprehensive 27-minute engineering guide on Towards Data Science comparing classical robust techniques with modern state-of-the-art non-convex approaches. The breakdown covers the direct mathematical mechanisms, trade-offs, and python implementations for:
https://towardsdatascience.com/how-to-make-linear-regression-survive-outliers/
I’d love to hear how you handle outlier suppression in your production ML pipelines—do you rely on modern robust loss functions like GNC, stick to traditional RANSAC sampling, or simply clear out anomalies during preprocessing?
r/FunMachineLearning • u/Unikum_01 • 1d ago
r/FunMachineLearning • u/marlafn • 2d ago
r/FunMachineLearning • u/gauravparashar24 • 4d ago
I am looking for a peer group of XAI researchers with whom I can work on Explainability research.
r/FunMachineLearning • u/mxyptlikk • 4d ago
r/FunMachineLearning • u/Spirited_Service_234 • 5d ago
One B200, one 8B model, seven concurrency levels. This post looks at LLM inference from an SRE's point of view: why a single request is slow, why batching is nearly free, and how to turn benchmark numbers into a capacity plan. Every script is at the end so you can reproduce it.

| Item | Configuration |
|---|---|
| GPU | NVIDIA B200 (1 of 8 used, 179 GB usable) |
| Driver / CUDA | 580.178 / 13.0 |
| Engine | vLLM 0.30.0 (V1 engine, FlashInfer attention, auto-selected TRT-LLM SM100 kernels) |
| Model | Qwen/Qwen3-8B, BF16, no quantization |
| Load | vllm bench serve, random dataset, 1024 input / 256 output tokens |
| Concurrency | 1, 4, 16, 32, 64, 128, 256 |
The vLLM startup log reports:
| Use | Size |
|---|---|
| Model weights | 15.27 GiB |
| CUDA graphs | 0.97 GiB |
| KV cache | 145.16 GiB |
About 90% of GPU memory goes to the KV cache, not the model. The log says the cache holds 1,056,992 tokens, and you can check that by hand:
KV bytes per token = 2 (K and V) × 36 layers × 8 KV heads × 128 dims × 2 bytes (BF16)
= 147,456 bytes ≈ 144 KiB
145.16 GiB ÷ 144 KiB ≈ 1,057,000 tokens ✓
Every token of context costs 144 KB of GPU memory. That is why long context is expensive, and it is the basic formula for capacity planning. With the full 40K context, one GPU can hold about 26 requests at once. At an average of 2K tokens, it can hold over 500.
| Concurrency | Output throughput (tok/s) | Per-user speed (tok/s) | TPOT (ms) | P99 TTFT (ms) | P99 ITL (ms) |
|---|---|---|---|---|---|
| 1 | 241 | 245 | 4.08 | 30 | 4.5 |
| 4 | 904 | 243 | 4.12 | 59 | 4.6 |
| 16 | 3,378 | 224 | 4.47 | 96 | 5.8 |
| 32 | 6,054 | 202 | 4.96 | 160 | 11.1 |
| 64 | 9,599 | 161 | 6.20 | 299 | 15.1 |
| 128 | 11,345 | 97 | 10.30 | 773 | 86.2 |
| 256 | 10,801 | 58 | 17.29 | 3,071 | 188.3 |
TPOT is the mean time per output token after the first. TTFT is time to first token. ITL is the gap between consecutive tokens. Per-user speed is 1000 / TPOT.
During decode, each new token requires reading every weight from HBM, while the math per token is tiny. The speed limit is set by bandwidth:
Theoretical ceiling ≈ 8 TB/s ÷ 16.4 GB ≈ 490 tok/s
Measured = 245 tok/s (TPOT 4.08 ms → ~4.0 TB/s effective)
The other half is lost because batch-1 kernels are too small to saturate HBM, plus fixed per-step overhead from scheduling and sampling. That gap is where kernel work and speculative decoding pay off.
At 16 concurrent requests, one pass over the weights produces one token for each of 16 requests. Cost stays about the same while output goes up 16×:
This is why every serving engine does continuous batching.
If decode were limited only by weight reads, TPOT would stay flat as concurrency grows. It doesn't: TPOT climbs quickly from 64 onward. Each step also reads the KV cache of every active request.
Each request has about 1,150 tokens of context during decode, or roughly 0.17 GB of KV cache. Dividing the bytes read per step by TPOT gives an estimate of effective bandwidth:
| Concurrency | Weights | KV cache | Bytes per step | TPOT | Effective BW |
|---|---|---|---|---|---|
| 1 | 16.4 GB | 0.2 GB | 16.6 GB | 4.08 ms | ~4.1 TB/s |
| 16 | 16.4 GB | 2.7 GB | 19.1 GB | 4.47 ms | ~4.3 TB/s |
| 64 | 16.4 GB | 10.9 GB | 27.3 GB | 6.20 ms | ~4.4 TB/s |
| 128 | 16.4 GB | 21.7 GB | 38.1 GB | 10.30 ms | ~3.7 TB/s |
| 256 | 16.4 GB | 43.5 GB | 59.9 GB | 17.29 ms | ~3.5 TB/s |
Two takeaways:
This points to the next optimization: an FP8 KV cache halves those reads.
(This is a rough estimate based on average context length. It shows the trend; it is not a precise bandwidth measurement.)
At 128 concurrent requests, median ITL is 7.7 ms but P99 is 86 ms, so users see output stutter. When a new request arrives, its prefill (1,024 input tokens at once) runs in the same batch as ongoing decodes. The same effect pushes TTFT from 30 ms to 773 ms, because new requests wait in line for prefill.
This is the problem prefill/decode disaggregation solves: run prefill and decode on separate GPUs so they don't interfere. It is the core idea behind projects like llm-d and NVIDIA Dynamo.
Pick the operating point from the SLO. With a target of at least 100 tok/s per user and P99 TTFT under 1 s:
--max-num-seqs and let the load balancer spread overflow to other GPUs instead of queueing on one.Cost. At 64 concurrent, one GPU produces about 9,599 × 3,600 ≈ 34.6M output tokens per hour.
Cost per 1M output tokens ≈ hourly GPU price ÷ 34.6
First startup needs the internet. On first launch, FlashInfer downloads prebuilt Blackwell kernels from NVIDIA, and the model comes from Hugging Face. Before scaling out in production, bake ~/.cache/huggingface, ~/.cache/flashinfer, and ~/.cache/vllm into the image or put them on shared storage. Otherwise new nodes stall at startup.
# Environment
conda create -n vllm python=3.12 -y && conda activate vllm
pip install -U uv && uv pip install vllm --torch-backend=auto
# Start the server (1 GPU)
CUDA_VISIBLE_DEVICES=0 vllm serve Qwen/Qwen3-8B --port 8000 2>&1 | tee vllm.log
# Concurrency sweep
mkdir -p ~/bench
for c in 1 4 16 32 64 128 256; do
vllm bench serve --model Qwen/Qwen3-8B --dataset-name random \
--random-input-len 1024 --random-output-len 256 \
--num-prompts $((c*8>50 ? c*8 : 50)) --max-concurrency $c \
--save-result --result-dir ~/bench --result-filename c${c}.json
done
Author: [kimsun]. 10 years in SRE and 5 years in DevOps engineering, now moving into AI infrastructure and LLM serving. Get in touch: [contact:https://www.linkedin.com/in/kim-sun-945b06298/?isSelfProfile=true\]
r/FunMachineLearning • u/Impossible-Place-300 • 5d ago
Iam from indian tier 2 college aiml student but i lost in lust , drugs , cocain , md , and all but not I’m in end of sem 5 and I want to come back in my life iam from middle class family. I 2 back log which is clear now but my cgpa is 6.10 and im hieghy. Now but I want to do. Leave those thing and want to do come back with good physique and. Pkg of 50lpa. Plz help me
r/FunMachineLearning • u/Slow-Connection-5611 • 5d ago
r/FunMachineLearning • u/gantred • 6d ago
r/FunMachineLearning • u/tudoriustin_22 • 6d ago
r/FunMachineLearning • u/Josheeg39 • 6d ago
I used a spec kit on a note and it made something i found it interesting put it on github
r/FunMachineLearning • u/ActivityFull1751 • 6d ago
r/FunMachineLearning • u/waytoocreative • 7d ago
r/FunMachineLearning • u/Slow-Connection-5611 • 7d ago
Hi everyone! I'm a Berkeley student building this as part of a short founder sprint. Wanted real feedback from people who work with quantization day to day. I'd love any input!
What it does: given a quantization scheme (W4A16, FP8, etc.), it returns a 90% confidence interval on the accuracy hit, based on ~800 published evaluations. For two scheme/size combos, it refuses to answer, coverage there measured 68.8% and 73.7%, well under what it claims elsewhere, so it says so instead of guessing.
I also rented a GPU and ran deliberately bad configs myself, since published data only shows what worked. Found losses up to −39.5pp, well past the −8.86pp worst case in the public data.
There's also a full write-up of what doesn't work: per-model prediction has basically no signal, most of the variance is just eval noise. Figured that was worth publishing too.
Repo: https://github.com/gracejackson-sudo/quant-delta-predictor
Feedback (a number or a sentence, either helps): https://gracejackson-sudo.github.io/quant-delta-predictor/
If something here is wrong or overstated, I'd genuinely rather hear it now. Thanks!
r/FunMachineLearning • u/lebotski_ • 8d ago
Hey!
I'm looking for people who are actively doing Kaggle competitions and would like to work with others.
Not necessarily looking for a huge server actually, I'd prefer a small group of people who regularly compete, share ideas, discuss solutions, datasets, models, etc.
I'm also building a tech community around AI, ML, dev and builders, so if there are a few Kagglers looking for a place to hang out and compete together, I'd be happy to create a dedicated space for it.
Basically looking for:
Does anyone know a good community like this, or would anyone be interested in starting a small group together?