It’s more economical at a compute level, but not at the developer level. The moment you start customizing your crawler to use protocol X for site Y your scale story collapses.
It's a good point, but in practice it depends on how easy those customizations are to implement / maintain, and how much money and effort you save. At some point the compute cost can disrupt even the nicest scale story.
I think the path forward is that websites offer one path for humans, and another for scrapers. But the huge catch is the path for scrapers must be _genuinely_ and _reliably_ the more economical and scalable path (either through something like PoW arms races, or through fear of litigation). Otherwise they will continue to ignore instructions and intrude on the human path.
The temptation to offer different / inferior / limited content to scrapers will be too strong, so such solutions are doomed to fail.
This would likely work as "Cloudflare SideChannel", a (hypothetical) Cloudflare product that would let scrapers download the pages that humans actually visit, as they are added to the CF cache. It wouldn't work for the non-Cloudflare part of the internet where humans connect directly to the servers that have their content.
The moment we establish a standard for offering a "optimized for scrapers" version of a site, people who do not want to be scraped will weaponise that to serve junk to scrapers... and scrapers will subsequently refuse to use it.
Largely because they're residential botnets in places like Brazil (a real example from one of my sites that was crawled to near-destruction). Someone could probably do something about this, but it's out of reach for individual site owners.
How good are you personally at iteratively and sequentially jumping through the legal systems of dozens of countries over the course of years, interleaved with genuinely difficult technical investigation, to unmask successive onion-layers of identity in order to unmask one offender?
Oh, and it also only takes a few minutes to reconfigure everything and invalidate those years of legal and investigatory work.
“Expect to lose some functionality, at least when accessing our resources anonymously.”
This seems fine to me. It would be a better world if we could have anonymous bulk data access. But if aggressive scrapers are bloating host costs, I’m fine with logging in.
Now, the flip side is that ONCE logged in, I want my bulk access. The worst of all worlds with when you demand authentication and then STILL block bulk access.
Case in point, I want to automatically download my Amazon and Target order records. This is easy to automate with playwright or whatever, but authentication stays annoying. My sessions expire quickly and I have to re-auth all the time. There should be an API to pull this data down.
This surprised me, so I looked it up. Basically YouTube pays ~55% and Nebula pays 50%.
This is complicated by Nebula's "creator owned" situation. Which is complicated [0]. It's something like:
- Nebula's assets live in Watch Nebula LLC
- Watch Nebula LLC is owned by 83% Standard Broadcast and 16% Curiosity Stream (CURI)
- Standard Broadcast "is owned by 44 total people, all of whom are creators"
So it's fair to say there are creators with significant equity stakes in Nebula. It's NOT fair to assume that the 50% of revenue left after share is distributed evenly among all creators on the platform via an equity mechanism.
Of course then there's Floatplane, which is entirely owned by LMG. At least a couple years ago they targeted "a 70/30 split" but arrive there via "a fixed cost menu for features [...] based on bandwidth/maintenance costs' [1]
I'm glad YT is still sharing a meaningful amount of money with creators.
Of cause, I assume it's difficult to compare 50% of paid subscriptions to 55% of ad revenue. I have heard that YouTubers earn per view from premium subscribers than from ad supported views.
percentages are meaningless in the ad space and I think topic specific CPM would be a more meaningful comparison. Even on youtube itself things like kids toy videos can make 4x or more per 1M views when compared to a Political show or something technical.
Sorry I missed your comment, Peter. I don't know that it's worth very much to describe. It's not sophisticated. Each component just has an attribute named data-ai-context and content that describes it. This needs no LLM. e.g. charts go into canvas, so their data-ai-context has a textual summary of the content (same as is fed to the charting library); buttons describe what they'd do. Then when you click the "Copy to AI", a piece of JS hits a backend to sign a 15 min token for the same user (as if they'd logged in but with a short token) and pulls all the data-ai-context attributes and gives you a prompt in your clipboard describing the site organization.
It just so happens that doing that tiny first bit allows even stupid models to navigate the site. After that, advanced models also sometimes query for the data-ai-context and they can mimic the Copy To AI to not have to nav the site.
e.g. content "Job Failure Rate: failed sync jobs as a share of jobs run, per connector per day, over 7 days. Day X: facebook=8 failed / 862 jobs; ga4=0 failed / 196 jobs;"
It's really just an alt text element and bizarrely worked the first time I tried it with old Opus 4.5 and I've kept it. Perhaps the biggest use is giving the agent the short-term token.
a first look through the comments on some first page threads and they seem far more thoughtful and considered. it generally seems like a nicer commenting environment that might actually do me some good moving forward.
https://en.wikipedia.org/wiki/.me
reply