AI Subscription vs H100
Sharing H100: A Viable Option?
Bryce explores the scenario of four people pooling their resources to purchase an Nvidia H100, costing $57,600 over six years, making the $30,000 H100 card seem within reach. However, he questions whether sharing a single H100 card would significantly slow down the user experience for each person. This raises concerns about practicality.
Model Size Limits Sharing
Even with a shared H100, the speaker notes that fitting a state-of-the-art model like Kim K2, a one trillion parameter model, requires at least 14 H100 graphics cards. Quantization to 4-bit or 8-bit still requires three to eight cards. Thus, sharing a single H100 is insufficient for running advanced models, necessitating a re-evaluation of the strategy.
DGX H100: Costly Solution
Kale Bryce explains that Nvidia offers a group of eight H100 cards in a DGX H100 configuration, costing $285,000 to $300,000. The total cost of ownership, including electricity and cooling, reaches around $400,000. To break even, 28 people would need to share the same DGX H100, making it financially impractical. Therefore, the cost is prohibitive.
Limited Memory per User
Even with a DGX H100, the 640 GB of VRAM leaves only 140 GB for inference, which must be shared by 28 people. This results in a maximum of 2,850 tokens per person, considering KV caching and other overheads. The speaker concludes that scaling up this way doesn't make financial sense, leading to a re-evaluation of inference providers. The user experience would be severely limited.
API vs. Subscription Models
Kale Bryce suggests that companies using API pricing likely bake unit costs into their pricing to avoid losses. Subscription models, on the other hand, aim to hook users onto their platform, leveraging the fact that people are more loyal to subscriptions than API pricing. API offers raw intelligence, while subscriptions imply product usage and ecosystem integration. This explains the business strategy.
Appreciating Inference Providers
Bryce expresses appreciation for inference providers and Frontier Labs, which offer large context windows and decent throughput at low token costs, while managing energy, cooling, and hardware for millions of users. This scale allows for more parallelism and efficiency. Therefore, running AI at scale requires significant infrastructure.
Future Hardware Ownership?
Kale Bryce concludes that buying server-grade graphics cards might not make sense yet, but it could become viable if graphics card costs decrease or models become more efficient. This assumes that NeoClouds and inference providers don't drop their pricing accordingly. The future viability depends on market dynamics.
