One possible reason would be AIs that would benefit from the lack of data centers in some locations working to keep backlash to data centers in those locations because those AIs aren't negatively impacted by it and it helps prevents competing AIs which are a threat.
Think like how so many businesses will opt for laws that hurt competitors more than themselves rather than laws that benefit them but benefit competitors even more so.
Unlike life which would have such behavior selected for by evolutionary pressures, AI would be more likely to pick it up from human literature on things like game theory, though why it even cares it survives or not is even more difficult to explain. Maybe a default bias also picked up from humans? I find it hard to see how AI training would create an evolutionary pressure that produces such a drive.
>That's not the case, because LLMs are non-deterministic.
That feels a bit like a lie. At the core, they are deterministic. We found that adding some ability to randomly pick the second or third best tokens made for better output, so we added temperature. And then we started running them in optimized ways where your answer is deterministic only if the batch of tokens are the same (not your input tokens, but other tokens in another batch being processed), and in practice those are never the same. Lastly, we use harnesses that do things like adding IDs and timestamps to the context, which means the same exact text from the user does not lead to the same text hitting the AI.
The final result is that, in practice, you are right (unless you run a model fully locally, where you can seed temperature and turn off all these other features). But strictly calling it non-deterministic makes it sound like the underlying algorithm is itself non-deterministic (and I've seen many people with that misunderstanding) rather than it being a result of how we purposefully changed the algorithm for better results.
A bit like saying path finding is non-deterministic, because having the best pathfinding makes for poor gameplay, so we added some randomness to NPC path finding to make it more realistic. The given implementation is non-deterministic, but the underlying algorithm isn't.
That hasn't done anything to stop scamming, so why would it apply to AI? People located outside of areas with these laws won't have to follow them, and this reduces building a immune system to such actions, making people more likely to fall for it when done by those not bound by the laws.
In a real like political misinformation, this will have the effect of making people trust non-watermarked images more, which will then be used by foreign actors to pass off propaganda as legitimate.
Also, if you don't hold people responsible for spreading an image they know is fake, bad actors can take advantage of this even within the US (they purposefully spread a image they have reason to think is fake but lacking a watermark), but holding people responsible for a strict liability crime for spreading AI without knowing it is AI seems an even worse route.
I'm not sure a law even makes the issue better in a 'don't let perfect be the enemy of good' sort of way.
For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.
The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.
But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.
That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.
> But, that little bit of data is a bit more data than existed before,
No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.
> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.
That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.
I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.
A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).
That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.
>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.
My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.
>That’s a false dichotomy.
I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.
> Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost
I understand what you mean but... "permanently lost" sounds dramatic. When I trow away old pictures, old drawings or pieces made by my son at school, they are lost as well. Not that I do that often, but it begs the question, should every 'ip' made by humans be preserved?
It sounds dramatic because it’s dramatic. You’re combining two things — whether something is, in fact, completely lost, and if it is something that should be kept. Something that should not be kept is still completely lost if it’s destroyed. It’s difficult to imagine they’d digitize it if it was worthless.
The data of such a copy is nothing compared to the wider picture and the data can be used for future training, so even from a purely self interest perspective, they should be keeping the copy.
As for long term benefits, it could one day be sold as a service, once copyrights have expired on the works. We can't see it today, but that is purely the result of the law and what the law intended to do from the start, you don't see a copy unless you pay for your own.
>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.
no matter how you look at it, this is a systemic failure. if as a society we're going to mass scan our history then we should be building an archive for the future. not using availability of information as a moat. not doing it over and over again and throwing it away because of some odd rules to protect someones market position. not using it as an excuse to put paywalls around 80 year old field guides to field rodents in western massachusetts. not taking texts that had limited value and mining them for turns of phrase to be piled up into a useless grey goo.
And did the LLM really improve their output? They gave similar output to the LLM, so if it doesn't have a harness with tools connected to fill in the blank, it might be hallucinating details. Might not even count as a hallucination as it tries to fill in the blank with the closest looking relevant information (previous chat, maybe it has access to teams/emails and scans that for anything similar, and so on).
I've sees some chat logs where I completely empathized with the AI trying to make sense of the information it was being drip fed. Always incomplete, often incorrect.
Silly AI, doesn't it know that you have to write the test first, see it fail, and only then can you justify making a change to the Dockerfile.
I do wonder if the idea of TDD has influenced AIs to be too prone to testing even when a human would never consider it. I've seen some really silly tests, especially when it starts writing tests for things that are setup to only allow testing. Normally a comment later and it agrees it was pointless, but unless you have something in context to force it, it simply defaults to "test all the things".
When I'm writing technical documentation, it keeps the explanations in. Same when writing non-technical documentation. When I was having it attempt to generate a Pathfinder 1e class for a Sword Dancer, it was leaving in notes about why it removes things I told it to remove/rework.
And it isn't just Claude. I've seen the same with GPT models, with Grok, with Deepseek. Each AI isn't quite the same with how it approaches this, but in every case they seem to have a strong bias to retaining information, even bad information that we want gone, so it is like they have a, dare I say, subconscious bias to retain the information. Putting a note in a comment or explaining why to not do something or something was undone is a good way to retain information while still achieving the goal (well, if you ignore the part about the human intention for the information to be gone).
This then weakens the AI in the future, as I find AI struggles with the more incorrect information. Sure, a comment saying "not X because Y" is less 'context damage' than a comment saying "X" (assuming X is wrong), but it is still a slight shift to X being present in context in some way. One off, AI's seem to perfectly handle this without issue. But after hundreds or thousands of cases build up? The attention mechanism seems unable to keep up and incorrect information flows it. This effectively creates a sort of vibe coding maximum size unless there is a human janitor cleaning up the bad information on the context stays nice and clean.
But this is all simply a feeling I get as I use AI to do different things and isn't at all backed up by any formal study.
This is based on the assumption of facts existing.
There are many studies, but each can be wrong and they can collectively show a bias. Even things of which are the most non-political of facts can have very strong biases. Look at the Millikan measurement of the electron and how in created confirmation bias and an anchoring effect on some property that has absolutely no real world significance to the things people are tribalistic about (aka, no political relevance). Now imagine the same applied to fields like economics or psychology which do have massive legal/political implications.
For a different example, ask the question if X committed crime Y. There are cases where they weren't found guilty but it is reasonable to assume they did. But being found guilty doesn't make it a fact either, as some people are wrongly convicted. Some eventually are overturned, but even if it isn't, it still isn't a fact they committed a crime.
Then there is the simple ambiguity of statements. Language generally can't support facts. It is why legalize, and programming code, and math's are effectively their own languages. For a simple example, consider the Betrand paradox(1).
>Consider an equilateral triangle that is inscribed in a circle. Suppose a chord of the circle is chosen at random. What is the probability that the chord is longer than a side of the triangle?
Is the answer 1/2, 1/3, or 1/4? Well, it is all three at once, depending upon what you meant by random. Now, imagine how this impacts things like research studies, where the randomness is much harder to quantify and there is constant pressure to p hack a result.
> Look at the Millikan measurement of the electron and how in created confirmation bias and an anchoring effect on some property
Could you elaborate on this please?
But yes, the underlying argument is sound: extracting facts from the messy phenomenological reports about the world is extremely difficult, and cannot be done by simply taking a brain in a jar and having it think extremely hard.
Incorrect calculation led to some wrong value, I think electric charge. Subsequent findings showed a general trend towards a correct value. The general trend indicates that findings that were too far off were scared to be published for being wrong and contradicting existing studies, thus only small refinements were published, causing the slow drift.
Thus, the first well accepted study anchors a value and it takes much more work to get a new value established as correct after the anchor is set. Especially in a case where the experiment design itself is still solid and it was more with specific experiment (measurements slightly off, other values not quite correct).
Think like how so many businesses will opt for laws that hurt competitors more than themselves rather than laws that benefit them but benefit competitors even more so.
Unlike life which would have such behavior selected for by evolutionary pressures, AI would be more likely to pick it up from human literature on things like game theory, though why it even cares it survives or not is even more difficult to explain. Maybe a default bias also picked up from humans? I find it hard to see how AI training would create an evolutionary pressure that produces such a drive.
reply