In many investment theses - like Nvidia's bet that demand for compute will keep growing - the first order assumption is usually correct. Yes, demand for more compute, chips, infrastructure is huge and each year some additional data centers will be built. Where such investment bets usually fail is in the second-order assumptions: Ie. the expectation of the growth of demand. This is where there's a high chance that the current expectations are likely exaggerated. So: demand is likely to persist for the foreseeable future but not increase every year. And that can upend the whole investment story. That can be enough to make these bonds a huge burden for Nvidia in the end. Not because people stopped buying more compute but because they stopped buying more every year.
What makes this insanely hard to predict is that the compute needed for the same quality output has roughly gone down 90% every 18 months for ~5 years.
1) We don't know how long that trend will continue, but you do know where to look for when it may end (if smaller sized models continue to compress the knowledge effectively of larger models).
2) We don't know when the appetite for higher cost models might go down and by how much if smaller models get "good enough" and price becomes far more important.
It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less.
It is also entirely possible that at some size - LLMs pick up some emergent capability that doesn't scale well to smaller sizes - and that there's an incredible boost to demand to get that capability.
I think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute; and we are not going to run out of economically useful things to do with it anytime soon on the demand side.
The harder thing to forecast for me is if we hit a wall on increasing efficiency, either on the model weights side or silicon side, with current approaches. If we have to switch to something like burning the model weights into silicon to continue to make gains, then the current math on general purpose accelerators might be upside down.
Nvidia has been playing a dangerous but profitable game since the Crypto boom.
but now I think they probably have bitten more than they can chew.
Apple already proved with their unified memory - that as long you have the capacity you can run capable models locally - thereby goes demand for inference if everyone is running some model locally.
For training - Chinese models have proved that you don't need the latest & greatest in Nvidia hardware. Same as TPUs.
More interesting take on Nvidia's position than I've come across before. One thing to be noted is 1) Nvidia is already making moves in robotics so even if their position in AI (moreso llms) diminished, they certainly have another big avenue arguably harder to just get into (although I'm not sure what efforts Google is doing for the tpu in robotics). Another point is Nvidia is still the main player in the west, that is, China certainly can and will create their own full stack without reliance on US companies. That puts Europe and other countries in an interesting, do you buy Nvidia because it's the only option or for security. That's to say I believe Nvidia's position relied on many different things being true at the same time, and we're moving towards an environment where those things are certainly being contested at (roughly) the same time.
Even in the west, Nvidia's dominance is bound to weaken. There is a notable uptick of articles on HN about people running large models on AMD hardware. And while I don't know official sales figures, I know we have trouble getting our AMD system delivered
AMD's software story is still a lot worse than Nvidia's. But patching up vllm to run one or two models you care about on AMD hardware is a much easier proposition than using them in most other fields of AI.
Nobody is saying they’re outright failing, but that they’re not going to be printing money the way they have been recently. Think about Intel circa 2010: most of their competitors like POWER or MIPS were marginalized, they owned the desktop and server markets with a bit of competition from AMD well contained, and their biggest desktop competitor (Apple) had just switched. A lot of analyst predictions … did not match what happened next. The same was true of Cisco a decade earlier. Both companies are still there but they don’t set the terms in their market segments.
I’m not predicting Nvidia will MBA themselves to death in the near future but I think there’s a tendency to overstate how profitable companies will stay. The more money Nvidia makes, the more motivated their competitors will be to get a piece of that market and the more customers will be looking for alternatives like the push into TPUs which the article discussed.
The current administration is definitely corrupt enough that you could imagine an anti-competitive deal of some sort but I don’t think there’s a way for even that to change matters because key competitors are well-connected American companies willing to play that game, too.
I think this is a great example of the disconnect people have in these types of conversations.
You can both become a company that supplies 80% of the world with your type of product, and then still have your stock go down in value.
All it takes is over evaluation by the stock market. Then a course correction from unsustained growth on growth (second order). So even if you continually replace YoY 80% of the world's hardware on a rotating business, but you don't increase market share or increase demand (aka growth)... Your business looks stagnant to the stock market, and there isn't really anything you can do about it. The best you can do is track inflation +/- 2%.
And that's why a lot of older established companies were dividend stocks. You don't expect to growth anymore, but that's not where the value is anymore... The value is in the reliable sales that will happen after infinitum because your company controls a majority share of the business... And that's ok! Unfortunately, silicon valley has created a philosophy of 'you gotta expand into new fields or your on the decline' - aka neo-monopolization
IMO focusing on the hyperscalers is kind of misleading.
Yes, for programmers and tech companies AI is kinda boring now, but AI integration in general is still kind of uncharted territory.
There are so many small companies and individuals just getting started with AI today and I believe a large the customer base (and revenue) is still untapped. Hell, I’m discovering new use cases regularly still and the average mismanaged 30 people whatever SaaS vendor probably didn’t even get started yet.
Phew. I thought we were just about to leave it, albeit not quite getting to space, just dissolving the financial system so our AI overlords can make some more harvester drones.
There’s a massive difference between “going up forever” and “point B is higher than point A.”
The current setup can’t sustain a downturn, even if yes 20 years from now point B is likely to be higher than present.
That’s the danger. Those that are going to get wiped out by the AI bubble burst aren’t wrong about AI being huge long term, they just put themselves in a position to not survive the storms that happen between points A and B.
He was referencing a book that made that case if you "squint". If you read that as actually a serious "this caused world war 1" statement rather than. This looks to have gotten some dominoes rolling that may have contributed to WW1 then that says more about you than it does about the article itself.
Ben is wrong; demand for compute, aka revenue backlogs, is mythical and will collapse, simply because of two reasons :
1. Circular investment/spending.
2. Too much capital in the system, so returns cannot be hit regardless because the barrier is too high. (Evidence being every capital cycle in history)
Ever free newsletter and talking head spouts narratives like this free. If you want something that quantifies and gives actionable information, you must do it yourself or pay for it. What are your below $2000/month sources for good analysis?
The website from this article. It's not free. Ben Thompson releases 1 free article a week but the other 3 articles published each week requires a $15/month subscription. Ben Thompson is also very influential in Silicon Valley and the overall tech/media industry.
I know him, I subscribed for a while. Even his paid content is lacking. I'm looking more Valens Research kind of analysis. SemiAnalysis is also good in the higher tier.
ps. Being influential in Silicon Valley just means you are influential, it does not mean substantial. Leopold is still influential and gets money thrown at him at $100s of million despite having no substance.
you might be looking for SemiAnalysis? I only read the free portions of articles but they have various paid options, mostly targeting investors with information and tools.
They are the marketing wing of the AI ecosystem. Their recent article on how SpaceX would drive 500B in data-center revenue was ludicrous-mode. Lets revisit this in a few years.
1) We don't know how long that trend will continue, but you do know where to look for when it may end (if smaller sized models continue to compress the knowledge effectively of larger models).
2) We don't know when the appetite for higher cost models might go down and by how much if smaller models get "good enough" and price becomes far more important.
It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less.
It is also entirely possible that at some size - LLMs pick up some emergent capability that doesn't scale well to smaller sizes - and that there's an incredible boost to demand to get that capability.
It's just very hard to predict.
The harder thing to forecast for me is if we hit a wall on increasing efficiency, either on the model weights side or silicon side, with current approaches. If we have to switch to something like burning the model weights into silicon to continue to make gains, then the current math on general purpose accelerators might be upside down.
> If we have to switch to something like burning the model weights into silicon to continue to make gains
I think that's already being considered semi-seriously [0][1]
[0] https://taalas.com/products/
[1] https://ir.amd.com/news-events/press-releases/detail/1296/am...
but now I think they probably have bitten more than they can chew.
Apple already proved with their unified memory - that as long you have the capacity you can run capable models locally - thereby goes demand for inference if everyone is running some model locally.
For training - Chinese models have proved that you don't need the latest & greatest in Nvidia hardware. Same as TPUs.
only time will tell.
AMD's software story is still a lot worse than Nvidia's. But patching up vllm to run one or two models you care about on AMD hardware is a much easier proposition than using them in most other fields of AI.
Trust me guys, it's over!
I’m not predicting Nvidia will MBA themselves to death in the near future but I think there’s a tendency to overstate how profitable companies will stay. The more money Nvidia makes, the more motivated their competitors will be to get a piece of that market and the more customers will be looking for alternatives like the push into TPUs which the article discussed.
The current administration is definitely corrupt enough that you could imagine an anti-competitive deal of some sort but I don’t think there’s a way for even that to change matters because key competitors are well-connected American companies willing to play that game, too.
You can both become a company that supplies 80% of the world with your type of product, and then still have your stock go down in value.
All it takes is over evaluation by the stock market. Then a course correction from unsustained growth on growth (second order). So even if you continually replace YoY 80% of the world's hardware on a rotating business, but you don't increase market share or increase demand (aka growth)... Your business looks stagnant to the stock market, and there isn't really anything you can do about it. The best you can do is track inflation +/- 2%.
And that's why a lot of older established companies were dividend stocks. You don't expect to growth anymore, but that's not where the value is anymore... The value is in the reliable sales that will happen after infinitum because your company controls a majority share of the business... And that's ok! Unfortunately, silicon valley has created a philosophy of 'you gotta expand into new fields or your on the decline' - aka neo-monopolization
https://deepmind.google/models/gemini-robotics/
Google is mostly the party behind the whole VLA principle.
Perhaps this has something to do with the economic dislocations and world wars between the 1870s and today?
Yes, for programmers and tech companies AI is kinda boring now, but AI integration in general is still kind of uncharted territory.
There are so many small companies and individuals just getting started with AI today and I believe a large the customer base (and revenue) is still untapped. Hell, I’m discovering new use cases regularly still and the average mismanaged 30 people whatever SaaS vendor probably didn’t even get started yet.
Building a business model on the belief that “this time is different” always finds storms on the horizon.
The current setup can’t sustain a downturn, even if yes 20 years from now point B is likely to be higher than present.
That’s the danger. Those that are going to get wiped out by the AI bubble burst aren’t wrong about AI being huge long term, they just put themselves in a position to not survive the storms that happen between points A and B.
1. Circular investment/spending.
2. Too much capital in the system, so returns cannot be hit regardless because the barrier is too high. (Evidence being every capital cycle in history)
Disappointed by the lack of Tom Cruise.
ps. Being influential in Silicon Valley just means you are influential, it does not mean substantial. Leopold is still influential and gets money thrown at him at $100s of million despite having no substance.
https://semianalysis.com/
They are the marketing wing of the AI ecosystem. Their recent article on how SpaceX would drive 500B in data-center revenue was ludicrous-mode. Lets revisit this in a few years.