Daily Shaarli

All links of one day in a single page.

April 3, 2025

Beyond Public Access in LLM Pre-Training Data: Non-public book content in OpenAI’s Models – Social Science Research Council (SSRC)
thumbnail

We just can't leave the topic of how the big model makers are building their training corpus unaddressed. This is both an ethics and economics problem. The creators of the content used to train such large models should be compensated in a way.

Between this, the crawlers they use and the ecological footprint of the data centers, there are so many negative externalities to those systems that law makers should have cease the topic a while ago. The paradox is that if nothing is done about it, the reckless behavior of the model makers will end up hurting them as well.

How crawlers impact the operations of the Wikimedia projects
thumbnail

Unsurprisingly, Wikimedia is also badly impacted by the LLM crawlers... That puts access to curated knowledge at risk if the trend continues.

The Fifth Kind of Optimisation

A good look back at parallelisation and multithreading as a mean to optimise. This is definitely a hard problem, and indeed got a bit easier with recent languages like Rust.

Minimal CSS-only blurry image placeholders

This is a very smart way to create pure CSS placeholders.

Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer

Could be interesting if it gets standardized. Maybe other forges than Gerrit will start leveraging the concept, this would improve the review experience greatly on those.

AI ambivalence
thumbnail

I somehow recognise myself in this piece. Not completely though, I disagree with some of the points... but we share some baggage so I recognize another fellow.