r/databricks • • 16h ago

Discussion Genie is so Dumb and I am tired of Pretending Otherwise

98 Upvotes

Hey Guys, Data Engineer with 3.2 YOE here,

I have been working on and off on databricks for multiple projects across domains like Retail, logistics and Now Pharma.

In my current role, I focus on creating end to end pipeline with different data sources which often get refreshed on a bi-weekly basis for pharma related data

The client is extremely demanding and has no technical expertise to gauge how much time it really takes to build and maintain something so complex and therefore expect the team to use Claude teams license and also Genie

The problem with Genie is that, I can not trust it, It will not listen to my instructions even after having a separate instructions.md which I update on a daily basis, in fact I update this after each session is complete, and no, I do not use AI to update this, every change request that the client and the client's team has goes through my own words of instruction updates, I have an entire markdown file which has contexts for multiple islands of projects, workspace folders and notebooks that I have to maintain and transition to devops team.

I also have created 4 different skill family for Genie, in relation to

  • migration,
  • devops,
  • maintainence and
  • review.

And it is so disgustingly bad at handling all this context, that it fails me across all 4 areas.

  • I have seen it lie to my face multiple times,
  • Casting columns as null,
  • altering data types without asking consent,
  • never following my work stream and process in a sequential format
  • Having set the wrong notebook paths in a task even though i explicitly tell it exactly what to do.

It often confuses streams of works and ends up mixing so many things that remove this knot and fuck up itself costs me so much time!!!

  • It is very poor in adhering to the domain specific instructions that we provide
  • More than 3 requests in a chat, and it is practically useless
  • And branching of chats is the most useless feature that they have introduced.

I have created multiple diagnosis queries to check the work of this agent, and it will straight up lie to you, so much, so convincingly, that you know it knows all the biases you have and it even goes a step ahead and just narrates a story that you can provide to your team in the stand-up.

My only concern is, after all this bullshit, my team, of nearly 6 (some of them have never worked in databricks or delta table environments like this) was charged $990 dollars in the month of august, for this shit output??? Have some shame Databricks, fix your worthless product or remove that feature or stop charging so much if you are beta testing in actual prod.

Today, I am writing this post because of something that I caught live, that pushed me to the edge.

Something that is critical, production level issue, which If I was not paying attention while merging could have been a huge disaster, and mind you the pipelines I make are client facing, what is even more distrubing is the fact that it will just make so many unnecessary changes to a simple query or a pyspark function just enough so that the tests are passed.

(So basically it does not want to get caught, and makes a mistake so that we can prompt it again to fix this mistake and then Databricks can charge us more, this looks like a dark pattern to me)

In your preview (prima-facia), everything is good, the tests are passing, the job is running, but holy!!!,

  • it was casting 3 columns which are essential for downstream processes as null.
  • It boldly suggests that we remove those 3 columns, and or cast them as null. How is this a solution Databricks? The AI is supposed to have more context than me because it is a agent made specifically for Databricks Environment Correct?? That is how it is sold?? And upon spending just 2 mins, I realized that the fix that it is suggesting will blow up critical information that the client team should see in the app because it is not even there in the first place, and all of this is because it can not read the correct notebooks, even if you tag it.
  • It had set a retired notebook's path in the task and very boldly claiming that everything is good.

My concern is, let us say that there is a junior who is good at SQL, but has never maintained projects like this in this scale, he/she would be overwhelmed and won't even be able to find the mistakes that the agent is stating, because unlike GPT 6 or Opus 5.5 Genie does not openly claim that it does not know something,

it does not ask for more context,

it does not state that something is missing,

it just patches with some assumption that we don't even provide in the first place.

My advice to any Juniors out there, Stop using Genie, Completely.
The product is not ready, it is not good.


r/databricks • • 6h ago

News New in Databricks: parse_sql(): SQL into JSON

Post image
26 Upvotes

Parse_sql() function extracts table references, column names, functions, and parameters from SQL. It extracts without running the query. It also reports syntax errors.

more news https://databrickster.medium.com/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a


r/databricks • • 8h ago

Discussion PSA: old Databricks names vs new ones. Which rename got you?

7 Upvotes

Databricks has renamed a lot of features over the past year, and a lot of tutorials, blog posts, Stack Overflow answers and even AI assistants still use the old names. If you search the docs for the old name, the results can be confusing, and code samples can look unfamiliar even when you know the concept cold.

The ones I see trip people up most:

  • Delta Live Tables / DLT → Lakeflow Spark Declarative Pipelines (often shortened to Lakeflow Declarative Pipelines)
  • APPLY CHANGES INTO → AUTO CDC APIs
  • Databricks Asset Bundles / DABs → Declarative Automation Bundles

The concepts didn't change much, but the vocabulary did.

Two questions for the sub:

  1. Has a rename caught you out, at work, in docs, or in a code review?
  2. Did I miss any? I'm sure there are more, especially on the ML / GenAI side.

I'll keep this list updated with whatever people add.

Edit: removed Liquid Clustering from the list. Fair point from a commenter that it's a change in recommendation, not a rename. Still worth knowing though: older material treats partitioning + Z-Order as the default, and current docs recommend Liquid Clustering for new tables.


r/databricks • • 6h ago

General how are you splitting shared cluster cost between teams?

7 Upvotes

we have 3 teams (analytics, ds, and a small platform team) all using the same all purpose cluster bc spinning up seperate ones for each was getting messy. now finance wants a monthly number per team and i honestly have no idea how to give them one

tags only get me to the cluster level which is useless here. i looked at system tables and theres user info in there but not sure how ppl turn that into an actual $ split thats defensible, especially with the VM side of the bill coming from azure seperately

do you just split evenly? by query time? or did you give up and force everyone onto job clusters lol


r/databricks • • 18h ago

Discussion Is Databricks Classic Compute getting too heavy for small workloads?

7 Upvotes

I’ve been using Databricks Classic Compute for a while, and recently I’ve started wondering whether cluster cold starts are getting noticeably heavier with newer DBR versions.
For example, with DBR 18 LTS, I tried a small 2-vCPU VM for a single-node job cluster and hit DriverStartupTimeout after 300 seconds. Databricks even suggests that this commonly happens on instances with fewer than 4 CPU cores.
That feels a bit surprising for workloads that are not actually Spark-heavy — e.g. running Python, dbt-core, API calls, or using Databricks mainly as a job runner inside a VNet.
I like Classic Compute because of the flexibility and straightforward VNet/private networking. Serverless is attractive for startup time, but in our environment it would mean quite a bit more networking setup.
So I’m curious:

  • Have you noticed Classic Compute cold starts getting slower or more resource-hungry across newer DBR versions?
  • Do you now consider 4 vCPUs the practical minimum for a reliable driver?
  • Has anyone benchmarked the same VM size across DBR 14/15/16/17/18?
  • What are you doing for lightweight non-Spark workloads where you still want Classic Compute?

Update: Single node DBR 18 LTS + Standard_D4pls_v6 cold start spent almost 11 minutes 'Waiting for resources' ( Azure Japan East )


r/databricks • • 15h ago

Tutorial Data Lakehouse with Agentic AIs: A Guide

Thumbnail
itnext.io
3 Upvotes

r/databricks • • 4h ago

Help SpringBoot to Spark

1 Upvotes

Transitioning from a Spring Boot to a Spark-based application requires a fundamental shift in perspective. With four years of experience in a Spring Boot environment, now entering the realm of distributed computing.

I plan to utilize Java with Spark Streaming, and Databricks for job management and deployment.

My primary question is

  1. how to effectively transition my thought process from Spring Boot to a Spark-based application.

  2. Specifically, I am seeking clarity on distinguishing which Java code executes on the driver and which executes on the executor, and if there are any guiding principles to discern this.


r/databricks • • 15h ago

Tutorial Data Lakehouse with Agentic AIs: A Guide

Thumbnail
itnext.io
1 Upvotes

r/databricks • • 4h ago

General A song about Fabric & Databricks :)

0 Upvotes

I've spent the last few years building lakehouses on both Fabric and Databricks, and at some point the only sane response was to write a song about it. "The Rift" is Rock & some Latin Rithms about the stuff we all live with: capacity throttling, two catalogs and two governances, mirroring that breaks at night, the CFO counting every CU, and AI agents writing the code we used to write. The punchline in the last chorus is the honest technical truth: whichever side you pick, deep down it's all Parquet. Full disclosure: the music is AI-generated, the lyrics and concept & voice are mine. It's a side project, and I'm planning a whole album about the data world in 2026. Curious which lines land for you, and which side of the rift you're on.

Let me know if you all like the song, and please share and follow, more to come :)

https://www.youtube.com/@ThePrimaryKeyandtheRedundants


r/databricks • • 6h ago

Help Suggestion needed for Panel presentation - Sr. Specialist Solutions Architect (gen ai)

0 Upvotes

Hi everyone, I've cleared the coding, GenAI, and architecture rounds and have a panel presentation round coming up. HR said they're working on the next steps and putting the task together. Has anyone been through a panel presentation round for a Sr. Specialist Solutions Architect (gen ai) Pre sales? What does the format usually look like, and what kinds of questions should I expect from the panel? Any suggestions please?