Egregoros

Signal feed

Timeline

Post

Remote status

Context

4
@sun

it's becoming apparent that we have like 5 classes of LLM haters

- the environmental fearmongering types who can't math

- the people who mistakenly think copyright is good for society and innovation and believe all output is somehow a stolen string of code

- the developers who are really into the compsci side and want all code to be perfect and concise (commendable, but unrealistic; these are the folks who should live in the university/research sector)

- the autistics who like the "puzzle" of chasing errors, diving into a debugger for hours on end and somehow derive joy from this

- the autistics who want to build a community around every git repo because their only friends are on the internet
@feld @sun
i hate llms because they've induced cloudflare-and-friends walls on every site—official-channels places i'm forced to use and tiny communities i actually care about—making web access far slower and less consistent

the number of places blocking tor has spiked, internet archive access in decline, opening tiny pages that used to be instant now has a 20 second delay every time, and i can no longer consistently access bank web portals, being filtered by something or other that can't be worked around with any combination of browser and vpn-or-lack-therof

also because search results for everything are now a thousand automatically acquired and populated domains of garbage
@ageha @sun this was all inevitable though, these scrapers are just what caused the tipping point.

This was always going to happen on a completely open network especially as bandwidth has increased making it easier to scrape, consume, attack, etc

It's our fault for not being forward-looking and taking the initiative to solve it

look at all the work being put into PQC to protect against attacks from theoretical computers that may never even exist -- this threat vector was identified, and encryption protects privacy and money and all sorts of things so people did something about it.

Nobody did anything about ensuring mass-scraping would be too expensive to do but not affect the experience of regular users
@ageha @sun

> mass scraping had no motivation before.

Only because we were ok with blessing a handful of companies (mostly search engines) with this power and expected them to not abuse it. But they often did. I've worked on projects that were absolutely getting crushed by badly behaving Apple/Microsoft/Google search indexing and had to respond while on-call to triage and block them.

Eventually we were going to lose trust in those search engines, and then what? How do we search? Well someone's gonna have to index it all over again. Storage gets cheaper, people start wanting to have their own private search engines... next thing you know you've got 10,000 indexers hitting your site.

It was inevitable. We just couldn't predict the catalyst.

Replies

1
@ageha @sun it would be better if instead of getting scraped/indexed that we moved to a culture of "we submit to YOU the content we want to be indexed" and then just build that functionality into websites/backends directly.

Maybe there could be relays like Nostr has and all the major search engines just subscribe and suck down all the content they want, filter/reject anything they don't want. But we control what's getting indexed because we are opting in and submitting what we want to be found.