Hacker Newsnew | past | comments | ask | show | jobs | submit | saithound's commentslogin

No. In the semiconductor industry, the "catch-up" player isn't normally spending less in absolute R&D terms.

Comparing the R&D costs of creating GPT-4o vs. DeepSeek V3 (the latest gen for which we already have good accurate numbers) it looks like the latter cost 1/20th as much to create.

If Samsung could catch up with TSMC for 1/20th of the cost, people definitely would say that TSMC has no moat.


Why do you think Chinese models cost 1/20th to train?

That's the ratio the widely published numbers give [1]. One does not have to believe the numbers [2], but those who do believe them are then justified to conclude that there's no moat.

Which numbers you believe is of course going to affect whether you think there's a moat or not. That's largely orthogonal to your TSMC/Samsung analogy I responded to. If you think the "moatists" are wrong because they believe the wrong numbers, that's fine, but then there's no need for the analogy.

[1] https://galileo.ai/blog/llm-model-training-cost

[2] https://medium.com/@theiand/how-can-deepseek-a-5-6-million-l...


But fundamentally, why is their cost 1/20 and is it sustainable in the next 10 years of competition?

Now that is a good and interesting question! Hopefully a "no-moatist" will share their reasoning.

Because they're distilling frontier models and that's a lot faster and cheaper than training a frontier model from scratch?

So why can't OpenAI/Anthropic also distill the good parts of free Chinese models? It's even better and easier for OpenAI and Anthropic. No poison pills as well.

Ultimately, that's what I need to be convinced. No one has put forth a good argument yet.

Clever architecture --> Ok but OpenAI/Anthropic can use these as well and they also have very smart people with their secret clever architectures

Distilling --> Ok but distilling means you will never be smarter than the original. Furthermore, reasoning is now hidden by private labs and they have poison pill answers for distilling if they can detect it. They will be able to detect distilling better and better.

Cheaper electricity --> Ok this is cancelled out by their chips being much less efficient due to not having ASML EUV machine access.

So I don't see why fundamentally their training costs are cheaper over the long term.

I'm looking for a no-moatist to convince me.


Labor. Smart labor would be much cheaper I'd reckon in China than in the US.

How much advantage in costs? What % of labor is training cost?

Mercor, Tacit Labs, Handshake AI... I suspect companies like these play a big part in model improvements, generating high quality benchmark/task-focused data for training.

However, these do require educated, white collar, workers.


Considering that frontier scientists and engineers in the US are currently taking home seven (or even eight, in some cases) figure salaries - pretty high, I'd reckon.

Would like to see the math since the claim is made.

I am not a "no-moatist" per se but one can argue their might be a plateau to how good a inference llm can become. If this is the case the playing field shifts to context, tools and harness, which are much cheaper to build an compete on.

Please just say what you want to say.

Not for long. Too late to get Business now to exploit this, since EH and the small credit-free Pro allowance will soon be restricted to Premium Seats ($100/m).

If we voice this opinion publicly, the most likely end result is that OpenAI will start billing our chat sessioms against our Codex budget too.

I thought they just did this? People were using some loophole to use their Chat sessions to power their Codex usage after their Codex quotas had run out.

No, it's still separate. I don't know what that exploit was though, so possibly they just patched that.

> Maybe it's because I'm American, but I can't imagine actually saying that to someone. I'd sorta expect them to throw hands if I talked like that about Jesus.

I mean, that's exactly what a "slur" is. They are not coined with the recipient's comfort in mind, and saying them can sometimes provoke exactly the sort of fight you imagine.

But yes, violent consequences ought to be particularly unsurprising to an American, since the US is, by Western standards, an exceptionally violent country. [1]

[1] https://pmc.ncbi.nlm.nih.gov/articles/PMC9535176/ (table 2)


> You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it.

If you have time, can you elaborate or give some examples of mathematical nuances?

I am evaluating Sol and Fable on a fairly large dataset of subtly flawed informal mathematical arguments (task is to identify and name propositions with substantially incorrect proofs in a larger body of text), and Sol is saturating the benchmark, while Fable is below 50% even with the most generous grading.

I don't work in the natural sciences, so I suspect you mean something different by "mathematical nuance".


Is this benchmark public? Anecdotally I have had decent results asking Sol to nitpick my proofs (mostly probability theory but nothing super dense). I have never tried Claude seriously, so I am very curious about what the failures look like with Fable.


Yes! All tasks are on github: https://github.com/harbor-framework/terminal-bench-science. They were contributed through PRs so the discussion and reviewing (before tasks were accepted) is also fully public.


> If it is known that A is provably true then one can study the consequences of A being true

But one can already study the consequences of P=NP right now. You don't need to know that it's provably true in order to do that.

Knowing an actual proof would be useful, but an oracle revealing merely that it's true (or even provable) without telling you the proof does not let you do anything you couldn't do before.


Some people (almost all mathematicians) wouldn’t want to spend time on consequences of a false statement. In the present discussion it’s not about letting me do something I can’t do now but about whether or not the endeavor is worthwhile.

A lot of people spent a lot of time and effort to prove or disprove the Jacobian Conjecture. AI solved it easily. It is increasingly becoming the case that humans are not as good at mathematics as computers. You are free to ignore computer generated proofs but I don’t think this position will win out in the long run.


> Some people (almost all mathematicians) wouldn’t want to spend time on consequences of a false statement.

No, people constantly prove statements of the form "if P=NP, then strange implication X". They do not consider it wasted effort at all, because of the contrapositive: if X is indeed very strange, they might be able to prove that it is false, and then they've settled P!=NP.


If a counterexample to a conjecture is found then all work toward proving consequences of the conjecture will cease. No one is trying to discover consequences of the Jacobian Conjecture now.

At some point an AI will prove a result that is so long and complicated that no human will understand it. This should not preclude people from using that result. In general, whenever the body of knowledge is increased it is a good thing. Even if it isn’t increased by humans.


Ah yes, Xorshift, the RANDU [1] of the 21st century [2].

There is no real use case for better non-CS generators, as explained by adrian_b back in 2021 [3].

[1] https://en.wikipedia.org/wiki/RANDU [2] https://arxiv.org/abs/1908.10020 [3] https://news.ycombinator.com/item?id=28886698


I’m not sure what point you’re trying to make, exactly, but a use case for better non-CS generators has always been stochastic simulation, especially simulation/sampling approaches that are bound by the number and quality of uniform variates per second.

As someone who has spent considerable time working in these areas, I still appreciate advances.


> I’m not sure what point you’re trying to make,

Have you skimmed the linked thread?

> especially simulation/sampling approaches that are bound by the number and quality of uniform variates per second

Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.

But even if you insist that you alone are doing some very weird stochastic simulation which is somehow bottlenecked on sourcing random numbers fast enough, the falling in planes phenomenon linked above would make xorshift-type generators a poor choice for most sorts of simulations. It introduces spatial correlations into any sort of lattice dynamics simulation (Ising model, percolation) and every high dimensional Monte Carlo integration. Beyond falling in the planes, since xorshift is linear over GF(2), it is also a particularly bad choice for nondeterministic cellular automata and Boolean dynamical systems which use parity, bit masks, or xors.

AES-CTR throughput on a modern CPU is higher than that of xoshiro256++, and much higher quality. No advances in non-CS PRNGs can beat that while maintaining the same quality. If your stochastic simulation is bottlenecked on random bits, CSPRNGs are still the way to go, and they don't interact in nasty ways with any dynamical system you can actually sinulate quickly.


> Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.

Actually, I spent a considerable amount of time in my doctorate and postdoc doing this.

Any kind of MCMC sampling of a simple model tends to be bound by the rate you can draw variates.

Examples of this include: Gillespie simulations of chemical kinetics, Ising and Potts lattice models (including their roughly bazillion variations), and anything resembling bootstrap or permutation sampling.

Just because your problems aren’t bound by the rate of drawing uniform variates doesn’t mean that these problems don’t exist. It just means that you have a narrow view.


I asked Vikash Mansinghka about this 15 years ago. He was using Xorshift as an RNG for a probabilistic inference on an FPGA. Why? It used few gates and was high (enough) quality.

Almost everyone should use a csprng, but iykyk.


Whether your recommendation is valid seems to be quite CPU-dependent. Cf:

  cpu: AMD Ryzen 5 5600X 6-Core Processor             
  BenchmarkAES_CBC-12             100000000               10.96 ns/op
  BenchmarkAES_CTR-12             83161234                14.36 ns/op
  BenchmarkPCG-12                 345336063                3.463 ns/op
  BenchmarkChaCha8-12             174143492                6.894 ns/op
  BenchmarkXoshiro256p-12         254343658                4.717 ns/op
  BenchmarkXoshiro256pp-12        266837442                4.496 ns/op
vs.

  cpu: Apple M4 Pro
  BenchmarkAES_CBC-14             162698020                7.370 ns/op
  BenchmarkAES_CTR-14             242501074                4.954 ns/op
  BenchmarkPCG-14                 197000988                6.083 ns/op
  BenchmarkChaCha8-14             237430095                5.050 ns/op
  BenchmarkXoshiro256p-14         252911710                4.738 ns/op
  BenchmarkXoshiro256pp-14        252656401                4.745 ns/op
Code: https://gist.github.com/kbolino/afbb86f3c9b2bd2f87272801d156...


You misrepresent the non cs ones by a factor of 5-10, even on cpu. First, you’re calling the prng twice per aes single call. That seems pretty dishonest already.

And using Golang? That ludicrously slow also.

To tell us what a cpu can do, do them using SIMD, you get pipelining and then many values per clock. Now try that with AES. Oh, you cannot, it’s not supported.

And people needing lots at full speed will do them on GPUs.

There is no world where even HW accelerated crypto comes close to non cs prngs on mainstream HPC systems.


This is a very weird hill to die on.

I do a lot of testing and designing of things like hash tables and filters, and having a really fast, non-CS generator is incredibly useful for being able to clearly identify performance bottlenecks in designs. PCG has been spectacularly useful for that purpose for me.


When was the last time a new PRNG helped you clearly identify a performance bottleneck?

As in, you were using state of the art generator X, and you couldn't see the performance bottleneck, but updating to a newer (faster, or same speed but higher quality) generator Y, and could subsequently identify the performance bottleneck?

If you're using PCG, not in the last 12 years.

(In a parallel comment I suggest trying AES-CTR for this use case)


It's not critical but if you gave me something that behaved statistically like PCG (i.e., I didn't fret about whether it was going to cause me weird problems) but was twice as fast I'd be happy and would shift to it - it would speed up profiling and measuring and that would be nice. We still find ourselves often pre-generating a list into memory to keep the prng entirely off of the measurement path. It wouldn't be magic, but I don't need magic. I like nice things that make my life a little easier in a small corner of my research. :)


Most modern RNGs should be faster than memory bandwidth (when optimized), so unless your list is small enough to fit in cache, its unclear if this is faster?


dgacmu: if you're writing C on x64, try AES-128-CTR (AES-NI, 8 way) using the header wmmintrin.h which has hardware accelerated primitives for this. An LLM can implement the RNG for you based on this comment if you want to test it out quickly. It should be faster than PCG, and higher quality.


Will do. I'm on vacation right now and losing my laptop for a few days, but seems worth trying. My recollection from the RNGs a decade ago (I'm dating myself) was that the AES approaches had higher latency but were quite decent, though slower than PCG. Curious how that's evolved.


> When was the last time a new PRNG helped you clearly identify a performance bottleneck?

While not a bottleneck as such, I contributed to a photorealistic path tracer using the Metropolis algorithm[1], and we got a 10-15% increase in samples/second when we switched from a decent to a much faster and better PRNG. Like you we didn't think the performance of it mattered much until we profiled it.

Granted this was a decade or so ago, would be interesting to compare the state of the art PRNGs.

Anyway, just pointing out that there can be real-world cases.

[1]: https://en.wikipedia.org/wiki/Metropolis_light_transport


As others here have pointed out, this is nonsense. The vast majority of PRNG calls on the planet are extremely high perf simulations, where crypto secure versions are a ludicrous cost in speed, energy, and sheer stupidity. That you and others do not understand is simply because you don’t see the places it’s required.

I’ve a PhD, have written papers on PRNGs, have worked in both cs prng and high perf prngs, have done decades of HPC projects, scientific sims. I get called in to develop precisely these high performance systems, and when you want to replace trillions to quadrillions of PRNG calls with one costing 10-1000x more, you’d get deservedly fired immediately.

You keep arguing about AES style code on a CPU. That’s not where people do high performance code. Try implementing AES and a fast prng on a GPU. You’ll soon find out how absolutely terrible cs-prngs are at performance. The measuremt isn’t how many ns per prng. It becomes how many thousands of prng generated per ns.

It’s bafflingly shortsighted for people with zero work in this area to continue to argue this. Choose the right tool for the job. Don’t project ignorance as knowledge. Both are useful advice.


Part of your argument was that there can't be an application for fast uniform pseudorandom numbers not just that xorshift by itself is not a very good PRNG(which I do agree with although it is an interesting sequence).

If Intel, AMD and Apple add a xoroshiro or PCG instruction and it produces pseudorandom numbers significantly faster than accelerated AES on those architectures, how does that affect your argument?

On the other hand, most simulations have moved on from random numbers to non-random space-filling sequences so the only non-CSPRNG application would be rolling fair dice for games and even there there is an argument to be made for CSPRNGs. So, perhaps I agree with you on the bottom line.


> no real use case

Yes there is. Not every system has the need or the resources to maintain a secure random sequence. You may also want a reproducible pseudo-random sequence in generative code that logs the seeds. Because of the misguided attitude that nobody needs these features, everyone who does need them has to roll their own now.


> Not every system has the need or the resources to maintain a secure random sequence.

I'm sure there's something, but that category has to be shrinking every year. What does such a system look like this decade, that needs random numbers but can't easily implement something like AES?

> You may also want a reproducible pseudo-random sequence in generative code that logs the seeds. Because of the misguided attitude that nobody needs these features, everyone who does need them has to roll their own now.

I don't know what difficulty you're referring to. Basically every CSPRNG can be seeded easily and you can log the seed.


I don't know about all that, but I use Marsaglias for generating noise samples in MCUs like Pico. It's the fastest option there is for such devices.


I agree! For many games (especially the ones running on old devices), xorshifts are pretty good! Not all applications need cryptographically safe generators! And sometimes it's ok to trade complexity for speed!

But I agree that if you're building something new aimed for modern devices, Marsaglia's xorshift128 wouldn't be my first choice! But I'd definitely want it in my RNG library for backward compatibility!


If you are interested in performant retro PRNGs, you can find a link to mine in my profile, if I am not mistaken. The counter is based on xorshift. As of the time of release it passed all tests that could be passed with only a 2^32-1 period.

On Z80 it was only about as fast as RC4 which is also a good option in terms of performance and quality but whereas RC4 has a huge state, this one only has a 32bit state, the rest staying in ROM.


> The article was showing the difference between mathematicians and engineers.

No. Many engineers AND mathematicians worked for a long time to get us to a stage where Amazon can solve a billion SMT problems a day. To contribute, all of them had to understand the theory this article calls overrated.


Some mathematicians certainly did, but there's a very large undercurrent in CS, as well as Mathematics more in general, of utter disinterest for applications as well as the idea that the more general a solution, the more "worthy" it is. That was really obvious from the words of the professor cited in the article.


I would to say I support that view. What is the purpose of modern science if not to discover truths you can apply universally? If you state: ‘this particular apple falls to the ground’ thats not a scientific discovery, there must be some general applicability.


There is an entire field of static analysis that is dedicated to practically solving undecidable problems.


Would you disagree that, given the choice, a solution to every problem is strictly better than a solution so only some problems?


No. Because for let’s say the halting problem the general case says it’s unsolvable, but each specific case is solvable.


1. Not all specific cases are solvable.

2. That the worst-case is very hard usually tells you that many instances will be hard (unless you discover an easy subclass), as is the case here. And when many "natural" instances are easy, that means that the problem is more interesting than perhaps previously thought, and it requires and receives more research, not less. If most instances are near the worst case, it means you know all there is to know about the problem; when they're not, it means there's more to study.


Can you give a specific finite program which is undecidable?


A program that enumerates all theorems in ZFC and stops when it proves a contradiction (e.g. true = false). Encoding a program that is equivalent to that directly as a Turing Machine in merely 748 states: https://www.scottaaronson.com/papers/bb.pdf (meaning that we cannot prove an upper bound on the 748th Busy-Beaver number, but there are probably even smaller undecidable TMs).

But my favourite example (shown here in Java) demonstrates the difficulty of analysing simple, realistic programs without necessarily being undecidable:

    long foo(long x) {
        if (x <= 2 || (x & 1) != 0)
            return 0;
        for (var i = x; i > 0; i--)
            if (bar(i) && bar(x - i))
                return i;
        throw new Error();
    }
    
    boolean bar(long x) {
        for (var i = x - 1; i >= 2; i--)
            for (var s = x; s >= 0; s -= i)
                if (s == 0)
                    return false;
        return true;
    }
Even in this case where even the input space is finite (and so everything here is definitely decidable), we simply don't yet know whether there is some x for which foo(x) throws, let alone if we made the input unbounded by using BigInteger instead of long.


Yes I disagree, because the algorithm to solve every problem takes longer to run than the remaining age of the universe.


Having read the article, I can confirm that it is.


While their math results are impressive, vibe coding their own web UIs with their subpar design models is really going to backfire if their plan is to attract new users with better free model offerings.

The Aug 6 update has forced the entry box to auto-format Markdown in an attempt to imitate Claude. The implementation is buggy and even simple copy-and-paste has gone entirely haywire. They also forgot to leave a switch to turn the confounded autoformatting thing off.

Chat mode in general is currently crawling with more UX bugs than a porch screen in summer.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: