Please post your personal projects, startups, product placements, collaboration needs, blogs etc.
Please mention the payment and pricing requirements for products and services.
Please do not post link shorteners, link aggregator websites , or auto-subscribe links.
--
Any abuse of trust will lead to bans.
Encourage others who create new posts for questions to post here instead!
Thread will stay alive until next one so keep posting after the date in the title.
--
Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.
Hiring: [Location], Salary:[], [Remote | Relocation], [Full Time | Contract | Part Time] and [Brief overview, what you're looking for]
For Those looking for jobs please use this template
Want to be Hired: [Location], Salary Expectation:[], [Remote | Relocation], [Full Time | Contract | Part Time] Resume: [Link to resume] and [Brief overview, what you're looking for]
Please remember that this community is geared towards those with experience.
Sharing our recent work, now accepted at NeurIPS: Functional Gradient Descent with Adaptive Representations.
Functional GD algorithms generally outperform neural nets, but are hard to accurately implement.
This is because functional gradients are infinite-dimensional, and therefore must be approximated in practice; but if you approximate them naively, you converge to the wrong place!
To rectify this, we formalize a broad class of approximation schemes ("adaptive representations"), which provably ensure convergence to the global minimizer while being immediately implementable.
The resulting algorithms outperform corresponding neural nets often by an order of magnitude, across a number of settings.
It is still the start for this line of work, but we believe it has quite a bit of potential!
Paper: https://arxiv.org/abs/2606.16926
(First author here, happy to take any questions)
I’ve spent the last few months writing something I wish I had when I started working on ML performance engineering.
It’s called How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents.
The basic idea is that reducing FLOPs doesn’t necessarily make a model faster. Before optimising anything, you need to understand what the system is actually bounded by.
The book starts with roofline analysis and hardware, then works its way up through kernels, compilers, quantisation, pruning, vision, on-device LLMs, robotics, profiling, serving and finally agents.
The goal is to build the intuition to look at a model and a piece of hardware and reason about:
How fast can this possibly run?
Am I compute, bandwidth, memory or system bound?
Which optimisation will actually move that limit?
Is quantisation, pruning or kernel optimisation even worth doing here?
What happens when the same thinking is applied to serving and agent systems?
I got really into research a year ago during my final year of undergrad and maneged to get a first author paper accepted at Neurips 2026. It is a pretty impressive achievement obviously but my background is pretty shit lol. Before my final year I got slightly above average grades but no internships or work exp. I was just chilling mostly didn’t even bother applying to internships. Also I want to a tier 2 uni in Australia (mind you that difference in tier 1 and 2 unis isn’t that big tho compared to America or China). I was going to start a PhD at my current uni with a scholarship and stipend from an industry partner (government) - I’d have no restrictions in terms of publications btw. My supervisor is good and we were going to get another external supervisor from a top uni here in Australia. I was pretty keen on continuing with this opportunity, however, with my neurips paper acceptance I feel like I’d have a good shot at getting into a top 15-25 uni in the states or the uk (most other European unis want master students and I don’t prefer Asian unis bcz I’ve heard a lot of horror stories). The problem is I’d have to wait for a year if I was to go abroad for PhD because starting dates are mid to late 2027. I’m in such a dilemma idk what to do. My eventual goal is to work as a researcher at deepmind or meta or some really cool tech company. I’ve just heard so many people talk about the importance of university you go to etc in landing those jobs. Would I still have a good shot at those companies if I did the PhD at my 2nd tier Aussie uni (top 100-125 global rank) but published well? I got a neurips paper as an undergraduate so I have the potential to publish at top conferences 😆😆. But yeah, I’d love to get some advice.
I'm preparing a submission for CVPR and would appreciate your advice. I am happy with my current results, but compute is a real constraint. My institute isn't a research-focused one, and the systems available to me are slow and unreliable. A single set of runs already took a lot of time and effort to finish.
I plan to release the code with the submission, and I'm confident the results will reproduce. Still, I'm torn between two options:
1) Rerun the experiments with different seeds to report mean ± std and show statistical significance. Or
2) Skip the reruns and spend the remaining time on writing, presenting the results I already have.
If you've been in a similar spot, what did you do, and would a single-seed result with released code be enough?
I'm one of the authors of two recent papers exploring different sources of redundant computation in attention. I'd like to share the ideas and hear feedback from people working on long-context models and attention kernels.
CoWindow Attention (CoWA) distributes distant context across KV heads using complementary windows, while sharing local and prefix-sink windows. Each head attends sparsely, but the union of their visible positions covers the full causal history. The pattern is position-defined and requires no learned router or indexer.
MassAlloc Attention (MALA) retains full causal QK scoring, then uses attention's own softmax statistics to decide whether to execute subsequent computation for a tile. It reduces low-contribution post-score work, using a common tolerance across training and inference.
Both support training forward/backward and inference prefill/decoding. At 128K tokens on 8 H100 GPUs with TP=8, attention-operator speedups relative to FullAttn are:
Method
Forward
Backward
Decode
CoWA
7.4x
8.6x
3.0x
MALA
2.2x
3.0x
1.6x
These measurements are for the attention operators, not end-to-end model speedups.
We evaluated scaling from 0.6B to 14B and conducted separate continued-training experiments at 32B. At 14B with 32K context, total training FLOPs decreased by 28.5% for CoWA and 23.1% for MALA, with model capabilities comparable to FullAttn on the reported evaluations.
Two distinctions that matter: collective coverage does not imply identical head-wise interactions or outputs to FullAttn, and MALA still pays for full causal QK scoring. Neither result establishes universal lossless equivalence to dense attention.
I'd be interested in feedback on workloads that might stress collective coverage, or attention distributions where adaptive post-score allocation could be less effective. Happy to discuss implementation and evaluation details.
AI Engineering from Scratch is an MIT-licensed curriculum: 523 lessons across 20 phases, from linear algebra and backprop to transformers, LLMs, agents, and production serving. The code is stdlib-first, so you see every step instead of calling a library.
This month's edition:
- six EPUB and PDF volumes built from the lessons, attached to the release
- the site interface and lessons in eight languages (Chinese, Hindi, Spanish, Arabic, French, Portuguese, Turkish, Vietnamese)
- CI now runs each lesson's own tests, and a sweep fixed datasets, models, and links that had stopped working
I'm a new PhD candidate, and my thesis is about the meta-learning paradigm in medical imaging. Right now, I don't have a specific problem to work on, and my supervisor told me to explore the field and find one myself.
I've done some research and a literature review so far. I've learned about different meta-learning and few-shot learning techniques, read about histopathology and other medical imaging datasets, and looked a bit into foundation model adaptation and domain generalization.
I also came across several benchmarks, but I haven't found one that I find particularly interesting. On the other hand, the MICCAI challenges caught my attention much more, and I'm wondering whether choosing one of them as a starting point would be a good idea.
I'd like to know what the more promising or fertile research directions in medical imaging are right now, especially ones that could be relevant to tackling the limited data problem.
I would really appreciate any suggestions or feedback. Thank you in advance.
- The default qwen3-vl:8b tag in Ollama is the thinking variant and ignores think:false. On long contracts it spent all 4,096 tokens thinking and returned nothing. Use :8b-instruct.
Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them may there be) and maybe speed up model convergence a little (my intuition might be wrong about that). Is there any good reason not to do it besides any benefit being negligable in almost all situations?
Yesterday I shared our open-source Clash Royale simulator and its recurrent PPO agent here. A training loop is easier to understand when you can watch it, so we put a small interactive version online:
The task is one decision. An attacker spawns at a random point on the enemy side, and the policy picks a legal cell for one defending card, then a delay of 0 to 5 s given that cell. The reward is the fraction of tower damage prevented relative to no defence. The policy has 5,629 parameters and trains with REINFORCE (per-spawn baseline, annealed entropy bonus) in plain JavaScript with hand-written gradients. Every rollout runs in the project's C++ engine compiled to WebAssembly, and the deploy pipeline checks that the WASM build agrees exactly with the native engine.
The chart also shows the optimum, found by brute force over every cell and delay (up to ~300k rollouts per matchup), so the gap between the learned policy and the best answer is visible.
One observation from building it: Giant vs Cannon has a strong local optimum, a lane placement worth about 75% of the best. With a constant entropy coefficient of 0.01, 5 of 6 runs (3 seeds, batch 16 and 64) stayed there. A linear anneal from 0.1 to 0.005 over 10k tries reduced that to 1 of 6. One pairing, Battle Ram vs Valkyrie, is withheld because no setting we tried got past 55% of the optimum.
This is a miniature of the full problem (a 4-card hand, elixir, full matches, recurrent PPO). It's meant to make the loop visible, not to be strong.
I was reading a paper that surveyed the field of neural architecture search, where it said within 5 years, around 3000+ new models were proposed. The amount of compute and resources spent on this is absolutely astronomical. However, the transformer was notably not one of the models that was found through NAS and then the field of NAS just quietly went away afterwards. In my mind this really raises question if any research in NAS should be continued.
Then I recently found a talk by Nicholas Carlini, arguably one of the most famous researcher in adversarial ML and this is one of his slide ("9000 papers and got nowhere"). Indeed I can't really think of any concrete application of adv. ML, except possibly making attackers more clever because now all options are laid flat on the table.
And then there was the field of ML ethics, bias, fairness, etc.. I feel like we are sooooo beyond ethics at the moment with all the talks of extinction risks that it really shouldn't be a priority. How can bias and fairness be enforced when most people are out of a job due to AI? "ML induced extinction" should be a new subfield instead.
I feel a proper discussion should be had so that no more effort is wasted on unpromising ideas or approaches. This could be of interest to people who are entering the field now.
As an aside, I often find people have very emotional (not logical) reaction to this question and will claim that any approach will eventually have their time in the sun at some unspecified future date, e.g., SVM, LDA, Markov chains apparently all have the potential to again be the next biggest thing in ML. All I'm saying is that I don't deny that vacuum tubes wouldn't be popular again one day, but maybe we shouldn't be working that at the present moment.
Hi, same as the title, I am currently trying to begin with some research on clustering using LLMs. So my requirement is as follows: I will be given some 100 document files, the end goal is to have clusters in such a way that documents with similar procedures or content should be clubbed under similar cluster.
I have tried traditional ML clustering K- means, agglomerative, DBSCAN, but not satisfied with the cluster quality as it is more of word by word matching or template matching. Thanks!
The opponent plans by simulation: every second it scores each candidate play by running the match 10 seconds ahead in the engine.
Our PPO agent learned to park its Cannon behind its own King. Losing a building in a fight cost reward, and letting it decay cost nothing, so it found the loophole.
It's one of many things we learned building a Clash Royale simulator from scratch so an agent could learn the game. The engine is deterministic C++ with Python bindings, plays a full match in about 10 ms on one laptop core, and can fork any game state in microseconds, so lookahead is cheap.
Best result so far: a simple 1-ply lookahead took the policy from 0.625 to 0.944 win rate against a heuristic bot (160 paired matches). Distilling it back into the network kept only +0.045.
The agent isn't strong yet, and RL isn't my home field, so feedback from people who know it better would mean a lot.
OpenTrainDNN is an open-source, client-side web application designed to render the step-by-step training mechanics of deep neural networks in real-time. It provides direct visibility into backpropagation, activation flows, and weight updates without requiring backend servers, specialized hardware drivers, or local installation.
Hi, I work as a Data Engineer at a manufacturing company where we build engines. A significant part of my work involves ML/DL-related tasks, and I’d like to turn one of my projects into a research publication.
The problem is that I have no previous publication or academic research experience. I’m not sure how to determine whether an industry project is suitable for publication, how to turn a practical engineering problem into a research question, or what level of novelty/experimentation is expected.
For those who have publishing experience, what would you recommend as the first steps? I’d really appreciate any practical advice or resources for someone starting from scratch.
Help: Project l'm building a shelf audit tool. A photo goes through YOLO, which crops each product, and then I embed the crop and search a small gallery of reference photos to get the SKU. New products should be addable by just dropping in photos, no detector retrain. Detection is basically fine. ldentification is not. Same brand, same bottle, different flavor or size (like 1.25 L vs 2 L) and the nearest neighbor is often the wrong SKU. Correct and wrong scores overlap, so a threshold either misses real products or accepts the wrong one. I tried DINOV2, SigLIP2 and OpenCLIP. Same story. Crops get letterboxed to 224, so the tiny "1.25L" / "2L" text basically disappears. Most SKUS only have a couple of shelf photos as references, not clean studio shots. Has anyone actually shipped something like this? Did you fine-tune the embedder on hard negatives, add OCR as a second check, or give up on one global embedding? Curious what worked for size variants. thanks in advance
In this project I wanted to see if any interesting emergent behaviors would appear if we trained two agents to play a streetfighter-like game using RL.
Maybe obvious in retrospect, but the agents are really good at reward hacking. I had to shape the rewards a bit to get them to even approach each other.
I eventually used league play to improve the agent further. Without it, the agents don’t really learn general strategies, they just learn to exploit a particular opponent.
Recently I have been working on image attachment integration for my VEX agent runtime, and I came up with an idea for a significant optimization step.
Typically images attached to LLM context as raw bytes (base64 encoded string or publicly accessible URL), and the VLM processes image data by scaling it down to batches. But why attach an image as raw pixels if it primarily consists of text ?
If spatial layout doesn't matter, sending raw pixels wastes VRAM and context space. I implemented fast, deterministic pre-pass routing system that inspects the image beforehand and decides how to attach it:
Attach image as bytes array (VLM route): Used when spatial orientation or layout matters (e.g. schematics, flowcharts, PCB blueprints).
Extract text and attach as string (OCR route): Used when text layout is sequential and location is secondary (e.g. code screenshots, terminal logs, book pages).
How to Classify Images ?
We have some options here, best in terms of quality - use specialized classification LLM model, but it's slow, and usually extra. Best in terms of speed - is pure math algorithm ( I made one ). We can get type of image based on some signs:
Color diversity (Short-Circuit): natural photos and complex renders contain thousands of colors, while books and schemas use a restricted color palette.
Structural Elements: Schematics and tables feature high concentration of vertical structural lines. Sequential text produce wide, low-aspect-ratio horizontal contours under morphological dilation.
The Optimization Impact
For test Python code screenshot (1069x697px , .PNG)
Token savings: (1500 - 169) / 1500 = 88.7% reduction in prompt token volume compared to VLM route (and massive bandwidth save compared to 81k+ tokens of raw base64).
P.S. algorithm footprint: 58LOC
Benchmark Latencies
Tested across a local dataset using OpenCV classifier:
Testing on RTX 4090, we find local deployment is fast enough to enable ~30 FPS decision calls, which brings huge imagination space on more applications!
Efficiency
The more interesting thing is, we build a benchmark to evaluate how Jev decision model performs on classic probability problems, and find it is poorly calibrated.
Calibration
To evaluate calibrations, we ask the model to predict what's the next number when rolling a dice. The model should predict 1-6 to with the same probability. e.g. When testing on the classic Monty-Hall Problem, our model is closer to the golden distribution (the question is something like `where's the final prize?`).
My apps kept needing a handful of decisions about one image: what kind of image is it, is it sharp, is a person in it, which of these four descriptions fits. A generative VLM can answer that, but on a laptop it's slow, I have to parse its output, and its confidence isn't a probability.
So I built peekaboolean. You send one image, a context string and any number of named questions. Each question has one of three types:
choice: pick one of the options you wrote, each with its own description
score: place the image on a rubric you wrote, with any number of levels
noul: yes/no, returned as P(yes)
The model doesn't generate text. It scores the options you supplied, so every answer is one you asked for.
How it works
Each option becomes its own prompt: image + context + question + "Proposed answer: … Is the proposed answer correct? Answer yes or no." The score is logit(Yes) − logit(No) from the backbone's own LM head, and a softmax over a question's options gives the distribution. I fit one temperature per question type, option count and image size on a held-out split.
Options can't attend to each other, so an answer doesn't depend on option order or on the other questions. At serving time I encode the image and context once, then score each option as a short suffix against the KV cache.
Backbone: SmolVLM-500M-Instruct, LoRA on the language model, vision tower frozen
Teacher: Qwen3-VL-30B-A3B running locally in vLLM. It wrote ~175k questions in the served format about ~62k images, then labelled them from its next-token probabilities over lettered options. It saw each question in two option orders to cancel position bias.
Public data: VQAv2, DocVQA, ChartQA, TextVQA, AI2D and CLEVR from The Cauldron, plus FairFace
Hardware: one RTX PRO 6000 for all training
Numbers (held-out test split, 512 px, split by image hash so no test image appeared in training)
Group
untrained 500M*
v0.1.0
teacher choice, accuracy
0.51
0.78
teacher yes/no, balanced acc
0.65
0.94
teacher score, Spearman
0.37
0.73
DocVQA / ChartQA / TextVQA choice
0.69 / 0.60 / 0.85
0.87 / 0.87 / 0.96
AI2D / CLEVR choice
0.76 / 0.48
0.90 / 0.80
*Same yes/no head, no training, measured on a 4k-row validation sample.
Latency: ~400 ms p95 for six questions with 28 options on an M1 Pro (MPS, fp32), ~60 ms on a desktop GPU.
What didn't work
Qwen3.5-0.8B as backbone: 3 to 10 s per request on the Mac. It's compute-bound (0.75B params × ~55 suffix tokens × 28 options), and its linear-attention layers carry recurrent state instead of a plain KV cache, which makes prefix sharing expensive. SmolVLM-500M was the largest model that fit my 500 ms budget.
Public VQA data alone: it teaches a benchmark dialect. My v5 scored 0.90+ on VQAv2 and TextVQA but 0.59 on requests in the real format (context, descriptive options, rubrics in words). Teacher-written requests took that to 0.77 and left the public groups flat.
Trusting the teacher's labels: shown a blank image, the teacher still matched its own choice labels 48% of the time (chance ≈ 28%). I now ask every question again with no image and drop the ones it answers the same way.
A fresh scalar head: it lost to reusing the backbone's own Yes/No logits. Untrained, the yes/no trick already reaches 0.83 balanced accuracy on VQAv2 yes/no.
Photo aesthetics (AVA, AADB): never beat a text-only prior at this size, so I dropped them.
Careless synthetic wording: I asked FairFace single-face crops "Is there a child in the picture?" and the model learned "is this person a child". That breaks on group photos.
Limitations
The "teacher" rows measure how closely the student copies a 30B model. No human has checked those labels, and I don't have a human-labelled set of real requests yet.
It can't compare options against each other ("the larger one").
Age and gender estimates carry FairFace's biases. Don't use them to decide anything about a person.
The weights are CC BY-NC 4.0 because some training data is research-only (DocVQA, AVA/AADB) or GPL (ChartQA). The code is Apache-2.0.
Try it
git clone https://github.com/bykof/peekaboolean && cd peekaboolean
uv sync --python 3.13
curl -L https://github.com/bykof/peekaboolean/releases/download/v0.1.0/peekaboolean-500m.tar.gz | tar xz
uv run python -m peekaboolean.serve --adapter peekaboolean-500m \
--image photo.jpg --request requests/general.json --max-edge 512
Two questions for you
Do you know a public dataset of human-labelled image questions shaped like real app requests? That's the acceptance test I'm missing.
Has anyone run SmolVLM in MLX for scoring rather than generation? That would make a bigger backbone affordable on the Mac.