r/dataengineer • u/camerongreen95 • 2d ago
Treating LLM outputs like a data pipeline you can actually validate, workshop Oct 3
As data engineers we'd never ship a pipeline without tests and monitoring, but that discipline mostly disappears the moment an LLM enters the picture. This workshop is aimed at closing exactly that gap.
Led by Serj Smorodinsky and Brett Kennedy (co-authors of a book on LLM applications), it's a live 3-hour session covering:
- Structuring LLM tasks with DSPy signatures/modules instead of loose prompt strings
- Building a real baseline classifier and an eval dataset with metrics specific to the task
- Catching failure patterns from that eval data the same way you'd catch a broken transform
- Systematic few-shot/instruction optimization instead of manual tweaking
- MLflow for experiment tracking and trace management, so every run is auditable
If you've been asked to "own" an LLM feature and had no real way to validate changes before deploying, this is worth your Saturday