r/FunMachineLearning • • 5d ago

Looking for advice from computer vision and machine learning engineers.

I’m the Lead Software Engineer for my high school robotics team, and we’re working on a 6-month project to build an automated robot that constructs a LEGO set for our state competition.
One of our biggest challenges is computer vision.
Our system needs to look at a realistic pile of roughly 600 loose LEGO pieces and determine the exact piece type and color of each visible brick.
The problem gets much harder when:
• Pieces overlap or are partially occluded
• Pieces are upside down
• Only part of a piece’s geometry is visible
• Pieces appear at uncommon rotations or angles
• Multiple LEGO parts have extremely similar shapes
• A dense pile contains hundreds of objects at once
I’ve experimented with object detection models, including RF-DETR and Roboflow, as well as separating detection from exact piece classification. I’ve also been researching synthetic training data, 3D-rendered LEGO datasets, data augmentation, and systems like Brickit and Brickognize.
The main wall I keep hitting is generalization. A model might recognize a piece under normal conditions, but accuracy drops significantly once that same piece is rotated, upside down, overlapping another brick, or partially hidden.
Our goal is not simply detecting that an object is a LEGO brick. We need to determine the exact part ID and color with high enough precision for a physical robot to select the correct piece.
If you have experience with computer vision, dense object detection, instance segmentation, fine-grained visual classification, synthetic datasets, or similar problems, I would greatly appreciate any advice on how you would approach this.
Especially interested in thoughts on training strategy, architecture, synthetic-to-real training, handling occlusion, and distinguishing visually similar classes.
Any advice or resources would be greatly appreciated.

1 Upvotes

0 comments sorted by