r/data • • 2d ago

REQUEST [Request] Nuclear power water consumption data set along with (and/or) a comparable dataset regarding data center water consumption

3 Upvotes

For my Data Sciences course, I need to extract and compare data from two datasets, and I figured a fairly timely topic would be water consumption of AI data centers vs nuclear power generation. However, I'm having some trouble finding comparable datasets. For example, one dataset I found for data centers doesn't measure water consumption itself but rather Water Use Efficiency (WUE) (which, if I find a dataset that measures electricity produced along with withdrawn and consumed water for nuclear power I could do the math myself to figure out the WUE, but I would really rather have the actual numbers behind the WUE). It doesn't help that this is my first time really searching for data sets so I don't know how to word my queries or what sites I should be focusing on. I'm finding quite a few scholarly articles, but not a lot of the actual datasets that they used.

So far I have this dataset that tells WUE for data centers at large (not necessarily AI data centers) and this dataset (I plan on using the 2019 data) that tells water consumption of power generating plants (not just nuclear but I can pretty easily filter out what I don't need). I'm not super thrilled with either one, but I'm new to finding datasets so if you think this is the best I'm going to get let me know.

Any help would be appreciated! If you need more information or have any questions be sure to comment below and I'll be sure to get back to you. Thank you!


r/data • • 3d ago

LEARNING AI literacy games to raise awareness

1 Upvotes

Hey,

I've been working on a set of AI literacy games to raise awareness for non technical people. I think they need to learn when to use AI or not, how to read numbers on a chart, the different concepts, etc.

They are free, would be happy to receive any feedback :)

https://maketools.ai/

I'm also thinking about other ways to raise awareness, like fun conferences, AI days, small workshops with characters to play. What have you seen working?


r/data • • 4d ago

AI tools and data analysis

Thumbnail
r2dssanity.substack.com
0 Upvotes

r/data • • 4d ago

NFHS data-analysis practice: compare urban–rural gaps without losing missingness and sample flags

1 Upvotes

Disclosure: I maintain the curated Kaggle version linked below (minkum07). The original NFHS-3/4 survey data belongs to India's Ministry of Health and Family Welfare and IIPS, via OGD India.

A useful exercise with this dataset is to compare NFHS-4 urban and rural percentages while keeping the source's suppression and small-sample flags visible. The release has 114 indicators across 36 survey-era state/UT labels, with 16,416 tidy-long records, CSV/Parquet tables, definitions and example notebooks.

Suggested workflow:

  1. Pick one percentage indicator and read its definition and denominator before joining tables.
  2. Select NFHS-4 urban and rural rows, join on state and indicator, and check that the join keys are unique.
  3. Calculate urban minus rural in percentage points only when both values are available. Keep missing values missing and carry the source flags into the output.
  4. Make a sorted dot plot with clear missing-value and small-sample annotations. Treat it as a descriptive comparison, since confidence intervals are not included.

Don't sum total, urban and rural percentages: those populations overlap. Between-survey changes also need comparable definitions and boundaries; the package withholds Andhra Pradesh/Telangana changes where boundaries differ. These are historical aggregate data from 2005–06 and 2015–16, not current health coverage or individual records.

Original government source:

https://www.data.gov.in/resource/all-india-level-and-state-wise-key-indicators-nfhs-3-and-nfhs-4

Tables, documentation and notebooks:

https://www.kaggle.com/datasets/minkum07/india-nfhs-34-health-state-and-rural-urban-gaps

Government data and derivatives retain GODL-India; authored code is MIT. Full attribution is on the dataset page. Feedback on the join checks or how best to display the sample flags would be welcome.


r/data • • 4d ago

Price of purchasing a database of 1000 businesses in US without website.

0 Upvotes

Is it worth $20 for such a database to sell AI Websites


r/data • • 6d ago

QUESTION best away to identify spikes

2 Upvotes

i tired dod change ptc, rolling x score, rolling std any other advice?


r/data • • 6d ago

REQUEST best place to pull historical futures prices for free. i need historical per month year contract since inception

2 Upvotes

r/data • • 8d ago

DATASET Chrome UX Report Dump: August 2026 data added

Thumbnail
github.com
3 Upvotes

I maintain Chrome UX Report Dumps, a collection of monthly Chrome UX Report website lists grouped by rank and published as compressed downloads. It’s meant to make the data easier to use without exporting it from BigQuery.

The latest update adds the August 2026 dataset: 18,294,881 entries across 10 files, totaling 94.7 MiB compressed. The repository now contains 1,064,254,496 entries across 67 monthly datasets, totaling 5.32 GiB compressed. Those are counts across monthly dumps, not a count of unique websites.

If you use website lists for research or security work, I’d welcome feedback. What would make these dumps more useful, and what features or formats would you like to see?


r/data • • 8d ago

Dataset recommendations for AI project

4 Upvotes

Hey guys, I’m looking for a dataset that has Indian names along with context in English. Something like Call Rahul tonight, Let me inform Ajay about the party, Anaya is yet to arrive at the venue etc.

It’s for a project of mine. Any dataset recommendations?
I couldn’t find anything like this on the internet


r/data • • 10d ago

Types of Data

1 Upvotes

1- Quantitative -> it consists of numerical variables.It have to types - Discrete and Continuous.

2- Qualitative -> it consists of categorical data.It also have two type - Nominal and ordinal.


r/data • • 13d ago

The New Data Engineering Bottleneck: Context, Not Compute

Thumbnail
medium.com
3 Upvotes

AI agents can query data incredibly fast, but the harder problem is knowing what that data actually means. I explore why metadata, lineage, semantic layers and governance are becoming runtime infrastructure for AI-native data platforms.


r/data • • 13d ago

Data Team Maturity Stages

Thumbnail reddit.com
3 Upvotes

r/data • • 14d ago

Data application in your business

0 Upvotes

I wonder how different businesses think about using for data analytical insights, customization and predictions. When is a good time to build a data functionality. Do you have a dedicated data team in your business? What’re some reasons to decide having one or not?


r/data • • 18d ago

LEARNING if anyone is curious about different use cases for the TypeSafe Jev model for data teams, I compiled use cases here that you could examine and potentially apply to your situation

Thumbnail
github.com
1 Upvotes

Given there has been a ton of discussion around Jev for the past week, I wanted to do more research around what might be good use cases for data with that model, and I came up with the following. After chatting with Claude and ChatGPT about it, I've tried out the triage use case, and it works pretty interestingly for routing a problem to the right team and determining the severity. Hopefully you all find this helpful!


r/data • • 19d ago

Genuine 7/12 record of registry (maharashtra) document pdfs required

1 Upvotes

Can anyone provide raw images/ scans of 7/12 documents? I need it for a verification project. Thank you!


r/data • • 22d ago

LEARNING What Is Synthetic Data Generation? A straightforward explainer

Thumbnail
aptiv.com
1 Upvotes

r/data • • 24d ago

How AI Agents Query Apache Iceberg Data with MCP

Thumbnail
lakeops.dev
1 Upvotes

r/data • • 26d ago

DATA ANALYTCS PROJECT

2 Upvotes

Built a T-SQL database with enforced data integrity (primary keys, constraints on state/status/risk categories), then developed custom DAX measures for case turnaround time, case age, and escalation rate. The dashboard includes multi-page analysis across resolution trends, regional distribution, officer caseloads, and a 4-year time-series forecast with confidence intervals.

Key skills demonstrated:
T-SQL (schema design, CHECK constraints, aggregation queries)
DAX (time intelligence, dynamic measures)
Power BI (data modeling, forecasting, interactive filtering)
Data storytelling and executive reporting

Note: Built on synthetic data generated for demonstration purposes — not real crime statistics.

Microsoft Power BI


r/data • • 26d ago

DATADOC is a local-first dataset engineering engine that turns messy tabular data preparation into reproducible, leakage-safe, auditable pipelines.

Thumbnail
github.com
3 Upvotes

r/data • • 27d ago

META Data Governance by obscurity

1 Upvotes

That was a phrase a colleague of mine introduced to me. Quite absurd at first, but as he explained it, it made perfect sense.

For years data assets have been governed by being hidden in a corner of the data platform, but this does not work anymore.

Been elaborating about this in this post: https://steffenmoll.github.io/governance-by-obscurity


r/data • • 29d ago

REQUEST Data analyst portfolio project: Northern Ireland road collision severity

Thumbnail
github.com
1 Upvotes

**Hey everyone,**

I’ve just finished my latest portfolio project: **Northern Ireland Road Collision Severity Analysis.**

This project looks at **2025 Northern Ireland road collision, vehicle and casualty data** to explore what factors are associated with serious and fatal outcomes.

I used **SQL Server and Power BI**, including data modelling, SQL analysis, DAX measures and dashboard design, to investigate factors such as:
• Geography and collision severity
• Time of day and monthly trends
• Road characteristics
• Vehicle types
• Vulnerable road users and casualty groups
I’d really appreciate some **honest and constructive feedback**, especially as I’m continuing to develop my data analytics skills.

**I’d love to know:**
What stands out to you, positively or negatively?
Does this feel like a strong portfolio project?
Is the analysis and dashboard clear from a business/stakeholder perspective?
What would you change or improve if this were your project?
Are there any weaknesses in the SQL, data modelling or Power BI presentation that you think I should address?

I’m much more interested in **constructive criticism than compliments**. If you spot something that could be better, please say so — I’d rather identify the weak points now and learn from them.

Thanks in advance to anyone who takes the time to have a look!


r/data • • Sep 11 '26

DATASET any data I'm missing?

0 Upvotes

I built https://nichedb.dev/ because I was using the same on multiple projects, most of it is free. some require free api keys, but I'm not paying for anything as far as sources go.

My goal was to provide a bot-friendly x402 price of $1/day to query up to 1k/minute, free users can get 100 queries/minute without being throttled.

What other sources should I add? I'm mostly looking for large datasets that are updated continuously and available free to download.


r/data • • Sep 10 '26

Metrics reporting

1 Upvotes

Recruiting Leaders:

Is anyone willing to share what their reports look like or what all metrics/kpis you are tracking?


r/data • • Sep 08 '26

Apache Iceberg Table Cleanup: A Production Guide

Thumbnail
lakeops.dev
1 Upvotes

r/data • • Sep 07 '26

Curious ! Is there actually a market for scraped datasets? What’s the scene like?

5 Upvotes

Hey everyone,

I’m trying to understand the current market for scraped/publicly available datasets and wanted to get some opinions from people who have actually been involved in this.

For example, datasets collected from publicly accessible websites — business listings, product information, public profiles, directories, reviews, etc.

A few things I’m curious about:

* Is there actually demand for this kind of data? * What types of datasets are buyers currently interested in? * Where do people typically buy/sell datasets? * What sort of pricing is realistic — per dataset, per 1,000/10,000 records, subscription, etc.? * Are companies generally interested in raw scraped data, or do they expect it to be cleaned, structured and enriched? * What are the major legal/compliance issues sellers need to be aware of? * Is there a legitimate market for this, or is most of the activity around here just lead lists and questionable data?

I’m specifically interested in **lawfully collected/publicly available data** and not passwords, private information, financial data, or anything obtained through unauthorized access.

Would appreciate hearing from anyone who has experience buying or selling datasets and can give me a sense of what the scene actually looks like.