-
7 votes
-
Consequences of advertising and enshittification on the Internet
48 votes -
AI comes to Playtime; Artifical companions, real risks
11 votes -
The Dealer's Tarot - Modern games to play with a tarot deck
23 votes -
Steam Machine and Frame verification slides from their GDC 2026 presentation
20 votes -
Economic ideas and policy implementation: Evidence from Malthusian training in British Indian bureaucracy
10 votes -
Script for the Superman movie released by James Gunn
12 votes -
Peter Watts on Margaret Atwood and the hierarchy of contempt (2003)
13 votes -
Amazon to allow EPUB and PDF downloads of DRM-free Kindle titles
36 votes -
Marian Heretic: Issues 1 & 2
4 votes -
Applying Chinese Wall Reverse Engineering to LLM Code Editing
8 votes -
Learning to Be Me (1990)
23 votes -
MotoGP information deck on Liberty Media investors sections
8 votes -
Which adhesive should I use?
24 votes -
NASA - Graphics Standards Manual (January, 1976)
14 votes -
How the 4% rule would have failed in the 1960s: Reflections on the folly of fixed rate withdrawals
18 votes -
Education Recovery Scorecard February 2025 report
5 votes -
Move over toasters: Doom is now playable inside a PDF
34 votes -
(PDF) Living happily ever after? The hidden health risks of Disney princesses.
16 votes -
Valve adds "Powered By SteamOS" branding for third party hardware
29 votes -
Power creep and the collapse of the Roman Republic
8 votes -
Paged Out! #5 - hacker zine release
7 votes -
Best solution to extract PDF data?
Hi folks-- To those more knowledgeable than I am: What would be the best local solution to extract numerical data from a batch of PDF file reports? The values I want are interspersed among word...
Hi folks--
To those more knowledgeable than I am:
What would be the best local solution to extract numerical data from a batch of PDF file reports? The values I want are interspersed among word processor formatted tables and irrelevant text. The text and table formatting are (nearly) identical across reports. The data I want vary across reports. The PDFs are not of images...I can select and copy text without OCR. I have thousands to process, and the data themselves are confidential (I have clearance) and cannot be shared. I can use Windows or Linux but no MacOS.
I am technically inclined, so I bashed my head against regular expressions just enough to use notepad++ to find and delete most of the irrelevant stuff and make a CSV, but it's a hacky, imprecise method and not nearly automated enough for batches. For reference, I don't code for a living or even as a hobby, but I use R and bash, am familiar with IDEs, and can follow pseudocode well enough to edit and use scripts.
Any thoughts? Thanks in advance!
24 votes -
US Senate investigation into Medicare and Medicaid insurance providers finds they are using "AI" to deny care
35 votes -
New experimental evidence shows lack of employment effects of guaranteed income
20 votes -
Grokking KOReader
25 votes -
Valve handbook for new employees — first edition
38 votes -
Evaluating the significance of San Lorenzo Village, a mid-20th century suburban community
4 votes -
Impacts Project
8 votes -
“Upload moderation” undermines end-to-endencryption: A statement from Meredith Whittaker, Signal president
28 votes -
Paper showcasing a simulation of gravitational waves produced by a warp drive
6 votes -
PayPal USD (PYUSD) on Solana
3 votes -
Zilog discontinues production of original Z80 processor after forty-eight years
28 votes -
webtoon-dl: a cli for downloading webtoons as pdfs
17 votes -
Yuzu, popular Nintendo Switch emulator, settles with Nintendo for $2.4m and halts development and distribution indefinitely
76 votes -
Sampling: What Nyquist didn’t say, and what to do about it
10 votes -
White House urges use of type safe and memory safe programming languages and hardware
38 votes -
The Great Automatic Grammatizator by Roald Dahl (1954)
21 votes -
Designing a Framework for Measuring Inference Generation and Communicative Efficacy During Board-game Play
16 votes -
Consider the lobster
32 votes -
Fighting climate change via personal banking
12 votes -
2023 GCHQ Christmas Challenge
15 votes -
Reddit moderators of r/law and r/scotus filed an amicus brief in US Supreme Court first amendment case Moody v NetChoice LLC
62 votes -
Denmark leads the Women Peace and Security Index 2023/24, scoring more than three times higher than Afghanistan at the bottom of the scale
14 votes -
Biking the goods: How North American cities can prepare for and promote large-scale adoption of cargo e-bikes
8 votes -
Intelligent traffic control with smart speed bumps
7 votes -
Douglas B. Lenat - The Ubiquity of Discovery
4 votes -
US study: Law abiding immigrants: the incarceration gap between immigrants and the US born 1850-2020
9 votes -
Towards understanding the design of games that aim to unify a player’s physical body and the virtual world
2 votes -
On being a c̵o̵m̵p̵u̵t̵e̵r̵ ̵s̵c̵i̵e̵n̵t̵i̵s̵t̵ human being in the time of collapse
12 votes