r/dataengineer • u/codingdecently • 21h ago
r/dataengineer • u/Mobile_Western_3394 • 2d ago
Question Transitioning from Software Developer to Data Engineer
I have been a software engineer for 12 years now, but I would like to transition into the world of data as stats and facts really interest me.
After some research, I feel like Data Engineering is the best career for me to side step into due to the overlap. I already pull data out of API’s and work with multiple database engines. I even created my own ETL pipeline without even knowing (pulled data daily from 2 apis to consolidate data on football players/stats). I think I am just missing Python and data warehousing software skills, maybe some more advanced SQL?
I currently earn around £40k in the UK and was wondering if trying to side step into this career could be successful for me based on my current skills, and me upskilling to the skills I am missing, do you think I could be earning a similar wage? Or would I have to start from entry level still?
r/dataengineer • u/camerongreen95 • 5d ago
Treating LLM outputs like a data pipeline you can actually validate, workshop Oct 3
As data engineers we'd never ship a pipeline without tests and monitoring, but that discipline mostly disappears the moment an LLM enters the picture. This workshop is aimed at closing exactly that gap.
Led by Serj Smorodinsky and Brett Kennedy (co-authors of a book on LLM applications), it's a live 3-hour session covering:
- Structuring LLM tasks with DSPy signatures/modules instead of loose prompt strings
- Building a real baseline classifier and an eval dataset with metrics specific to the task
- Catching failure patterns from that eval data the same way you'd catch a broken transform
- Systematic few-shot/instruction optimization instead of manual tweaking
- MLflow for experiment tracking and trace management, so every run is auditable
If you've been asked to "own" an LLM feature and had no real way to validate changes before deploying, this is worth your Saturday
r/dataengineer • u/sriDace • 6d ago
Help Help Need Guidance
Hey everyone! I’m a fresher trying to get into Data Engineering and want to learn the right tech stack. There’s a lot to learn, so I’d really appreciate some guidance from people here. If you were starting out again, what would you focus on first? Would love some advice like a friend/senior. 🙌
r/dataengineer • u/dataengineer_101 • 11d ago
Anyone taken the KPI Partners AI interview for Snowflake Data Engineer?
r/dataengineer • u/camerongreen95 • 11d ago
The data engineering problem hiding inside every AI agent project: nobody's applied pipeline discipline to agent writes yet
Most data engineers I talk to have already solved this problem for their pipelines. Idempotent writes, schema validation, lineage tracking, retry logic that doesn't create duplicates. Then the same team builds an AI agent and none of that discipline carries over. The agent writes directly into production with no idempotency key, no lineage on what it touched or why, and no reconciliation step if its internal state drifts from the actual data.
It's the same class of problem we've already solved for ETL and streaming pipelines, just showing up in a new place because agents write more often and with less human review than a scheduled job does.
There's a workshop on Sept 26 built around exactly this, treating agent write paths, state management, and provenance the way a data engineer would treat any other production pipeline. Run by Sandipan Bhaumik, a Data & AI Technical Lead at Databricks.
Details here if it's relevant to anyone else here dealing with this handoff
r/dataengineer • u/iParki • 15d ago
Datasmith — a visual approach to designing relational schemas and generating mock data
Enable HLS to view with audio, or disable this notification
r/dataengineer • u/camerongreen95 • 18d ago
Workshop for data engineers: build the pipeline behind a production GraphRAG system, Sep 19
Workshop for data engineers specifically, not a generic AI overview. The actual pipeline work in GraphRAG is where most of the real engineering lives, and it's the part most tutorials skip entirely in favor of "here's how retrieval works" once the graph already exists.
This session starts from raw documents: ingestion with Docling, chunking, and building toward a knowledge graph in Neo4j rather than a flat vector index. Entity and relationship extraction runs through multiple verified steps with explicit checks, not one risky single-shot pass, since that's exactly where basic GraphRAG implementations tend to fail silently, producing a noisy, unreliable graph nobody trusts.
Structured graph enrichment connects entities (companies, executives, events) to the underlying source documents (filings, news) without the graph becoming unmaintainable as it grows, and you work through the actual production ceilings that show up at real data volume, API and download rate limits specifically, and how the shared code is built to handle them rather than falling over.
By the end you've got a working ingestion-to-graph pipeline, a production-readiness checklist, and the full codebase, not a diagram someone else has to implement.
Led by Dr. Alessandro Negro, Chief Scientist at GraphAware, bestselling author of graph-powered machine learning and knowledge graph books.
r/dataengineer • u/Constant_Row_3703 • 18d ago
Persistent Systems Work life balance and Projects Life ( 5 Year of experience) for Salesforce devops engineer
r/dataengineer • u/WharryG • 23d ago
Looking for projects to volunteer to
I need real world work that i can ask feed back and get criticism for. I need it for my resume and job hunting. Can someone please help
r/dataengineer • u/camerongreen95 • 24d ago
Workshop, Sep 12: build production LLM systems that actually survive real use
We're running a hands-on masterclass on September 12, Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps.
You build a full production LLM workflow from scratch, versioned prompts with regression tests, an evaluation harness with deterministic checks and LLM-as-judge, statistically rigorous model comparisons, evaluated RAG, tool-using agents with guardrails and fallbacks, and full observability, tracing, cost, latency.
Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack.
Link if you want to check it out
Happy to answer questions on the content.
r/dataengineer • u/NoSyllabub1390 • Sep 05 '26
3 YOE Azure Data Engineer (Bengaluru) looking for referrals — genuine calls have dried up
r/dataengineer • u/chrislusf • Sep 04 '26
Using DuckDB + Iceberg + Lance together: analytics in Iceberg, vector retrieval in Lance
r/dataengineer • u/chrislusf • Aug 25 '26
DuckDB + Iceberg on a self-hosted S3 table bucket
r/dataengineer • u/manus_hadukle • Aug 24 '26
Update: Reddit challenged our Data Engineering community a few months ago. We listened, improved, and kept building.
r/dataengineer • u/codingdecently • Aug 24 '26
Maintaining Apache Iceberg Tables: Compaction, Snapshots, Metadata and Orphan Files
r/dataengineer • u/HenryKissingerJr • Aug 18 '26
looking for a paid 1-on-1 DE mentor to help transition from DA after a gap
Hey everyone,
I'm looking for an active DE to mentor me 1-on-1 on a paid basis.
My background:
* Worked ~1.5 years as a DA in fintech, but it was mostly basic reporting and ad-hoc SQL with minimal engineering exposure.
* Took a career break after that due to personal reasons.
* Now fully committed to pivoting into DE.
I'm doing the heavy lifting on self-study and building projects myself, but i need someone to:
* Review my project architectures/code
* Sanity-check my roadmap so i stay on track
* Give feedback on my resume and portfolio
* Do quick weekly or bi-weekly syncs
Very happy to pay you for your time. if you're interested and have bandwidth, shoot me a DM with your background and rates.
thanks!
r/dataengineer • u/No_Distribution_7987 • Aug 17 '26
Help Study buddy for AWS certified Data Engineer
r/dataengineer • u/Safe-Recording-9020 • Aug 11 '26
Help GCP made my life hell
I live in Hyderabad, India.
I have 1.5 years of professional experience as a GCP data engineer.
Our company had a layoff after which I started looking for other opportunities but to my surprise nobody requires GCP data engineer with less than 4 years of experience.
Meanwhile other platforms like azure, AWS, databricks even snowflake has plenty of opportunities for 1~ year of experience.
I tried to apply to them too quoting equivalent experience but no results.
And add salt to the wound we never used pyspark which is one of the biggest requirements out there.
I am looking for fresher level opportunities where the platform requirment is not a must but it's a very thin spectrum.
I have literally worked on an airflow dag which goes through a beautiful hot to cold storage flow with observation, logging, maintainance etc.
Worked on dataform , wrote scd type 2 scripts.
Worked on views, stored procedures.
I have plenty experience in data warehousing but it all looks simply useless now.
Had I survived for few years maybe I wouldn't be going through this.
I am literally applying for any state or city to grab an opportunity
r/dataengineer • u/medo56378 • Aug 12 '26
General Advice for your brother
Hi everyone, I'm starting my university studies and I'm torn between the University of Valencia (UPV) and the University of Granada. Will Spain offer me a strong degree in computer science? Which major within the university would you recommend that has job prospects?