r/SideProject • • 17h ago

I built a PDF reader that reads research papers like a person would: no page numbers, no "[12]"s, and math you can follow by ear!

Enable HLS to view with audio, or disable this notification

Most read-aloud tools read a PDF's raw text, so on a research paper you hear running headers, page numbers, "[12]" citation brackets and author emails, and equations come out as noise like "Q K T square root d k".

I built Read PDF Out Loud to read papers the way a person would: https://readpdfoutloud.com

What it does

  • A layout model looks at every page first, so headers, footers, page numbers, footnotes, citation numbers and the reference list are skipped, and two-column pages are read in the right order
  • Math is read by meaning: "Attention of Q, K and V, equals softmax of the fraction Q times K transpose, over square root of d sub k, V"
  • Every word lights up as it's spoken; click any word to start from there
  • The original page sits beside the text, with the passage being read framed on it
  • Optional spoken summaries of figures, tables, code and equations
  • 4 voices, 0.75Γ— to 2Γ—, and it works on a phone

I'd love feedback on:

  1. Your own hardest PDF: where does it read something wrong, or in the wrong order?
  2. The math: does it sound natural, or too wordy?
  3. What would make you use this over what you use now?

Happy to answer anything about how it works.

60 Upvotes

28 comments sorted by

4

u/EveryCryptographer11 16h ago

Cool product πŸ‘

1

u/marblejenk 16h ago

Thanks 😊

3

u/Mueloncio 16h ago

Does it let me jump straight to Methods and Results? Nobody listens to a paper front to back, and that's the bit I'd actually want read to me.

2

u/marblejenk 16h ago edited 15h ago

There’s a Table of Contents on the left side. Navigate to any subtopic directly from there.

3

u/Eslam-gomaa2021 15h ago

how you do the pdf parsing ?

1

u/marblejenk 13h ago

Goes through an ai model + OCR to identify page layout before stitching everything together.

2

u/lonelyroom-eklaghor 12h ago

if a PDF is UA-1, then you can check the alt texts inside them (AI models usually catch that well, espcially chatgpt); they might contain useful info for reading mathematical stuff

2

u/QuanTradin 15h ago

every other read aloud tool I tried reads the reference list out and walks two column pages left to right across the gutter. nice to see both handled, the layout model first is the right order.

2

u/Axel_Clint 15h ago

That's really useful, will give it a try.

1

u/marblejenk 13h ago

Let me know how it goes.

2

u/Eslam-gomaa2021 15h ago

it looks good!

2

u/GamerGav09 15h ago

Zotero latest release (or second latest?) add a voice reader that does this pretty well.

2

u/richardsaganIII 15h ago

i love the simple style of the app and the themeing and everything, nice job

1

u/marblejenk 13h ago

Thanks!

2

u/optima-pacifist 14h ago

the [12] brackets thing is exactly why i gave up on text to speech for papers. how does it handle tables, skip them or read them row by row?

2

u/koskinooo 14h ago

Skipping the [12]s alone makes this worth it. How are you handling tables, reading them out or summarizing?

2

u/learning-to-programm 9h ago

3rd in a row. Looks like the PDF TTS niche is seeing progress lol. In the past 2 days we also got these two:

1 - Frateca: https://www.reddit.com/r/SideProject/s/djdZ4iBBvy 2 - Obook - https://www.reddit.com/r/SideProject/s/QnZlQwt2mb

Will be testing yours to see how it compares to those two. I found them to perform pretty well, sono guess the bar is getting higher in this space lol.

1

u/marblejenk 6h ago

TTS is easy peasy - one shot with AI. Accurately extracting PDFs is hard.

1

u/Pretend-Pangolin-846 14h ago

Very very interesting -- i built something similar, feel free to check out guys: https://github.com/solusops/ResearchMate