Edit: Now I think about it, this might be the cheapest way to run the DeepSeek V4 Flash 0731 on a dedicated inference server at original weights. I haven’t run mixed load benchmarks but I guess it’s possible to generate $3-$4 worth of tokens per hour and still maintain a usable per-user throughput.
WASDx 11 hours ago [-]
At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.
zhoutong 7 hours ago [-]
There’s also input tokens. For many agentic use cases input/output ratio can be 2:1 or even 4:1. Non-quantised DeepSeek V4 Flash costs $0.14 / $0.28 on most inference providers with ZDR.
When self-hosting the model, cache hits are basically free. RAM cache also helps with hit rate (prefixes may be around for hours instead of minutes).
To be clear this project doesn’t aim to achieve the best inference economics per token. MI300X doesn’t have native MXFP4 so it’s not even the right platform for the model. That’s why very few deployment recipes are available.
It’s interesting to me because MI300X is quite accessible to a small team with budget for just 1-2 GPUs. DeepSeek V4 Flash otherwise wouldn’t even fit on 2x H100s.
We can run several coding agents during the day and batch inference jobs overnight and serve the entire team with guaranteed privacy, without compromising precision or speed.
In fact we found that many inference providers are quantising the weights or even KV cache, and due to the low prices they serve at massive batches, resulting in unstable throughput. I ran GSM8K as a quick validation test and this deployment is “better” than the OpenRouter endpoint in a statistically significant way (I wouldn’t name the provider here). I will run some follow up benchmarks and update the repo when I find some time.
lnenad 10 hours ago [-]
830t/s is burst aggregate. ~500 is sustained and it's for 8 concurrent users. Meaning for $1.99/hour if you serve 8 users it's 8*$0.54, not just $0.54.
You shouldn't rent one out if you're just serving it for yourself, but from a financial standpoint if you sell to users you can take a 100% margin.
minraws 10 hours ago [-]
Bro 500 Aggregate. so that's 500 * 60 * 60 = 1.8M output which is .5$ at best... Not including pre-fill and stuff.
This is not the real margins, even if you are selling to 8 users it's 90 tps per median stream. So assuming that .6-.7$
This is not even remotely worth it.
You need to 3x this tps(~1500 tps) to be worth it, and that's what most providers are doing, at 20-30 users at 50-60 tps with better optimized batch processing and kernels you can make some profit.
lnenad 9 hours ago [-]
You are right, it's not 500x8 it's 90x8.
Almondsetat 10 hours ago [-]
You get privacy for 4 times the cost
krisknez 10 hours ago [-]
How is that economically viable? They are selling at a loss?
gpugreg 10 hours ago [-]
Agentic workloads are somewhere around 1%/0.5%/98.5% input/output/cached tokens. Cached tokens are pretty much free for inference providers (if they implement sparse and compressed attention properly) and throughput for input tokens is much higher.
Lets assume that you've got 2 million input tokens, 1 million output tokens and 98.5 million cached tokens to process. That would cost 2 * $0.14 + 1 * $0.28 + 98.5 * $0.0028 = $0.8358 with DeepSeek API pricing.
For comparison, it would take 2M / 8000 + 1M / 800 = 1500 seconds to process this amount of tokens with the linked framework, which is about $0.83 when we assume $2/hr for one MI300X.
However, other inference providers have 10 times higher prices for cached tokens, which results in a comfortable margin.
And we should not discount that DeepSeek also gets paid in data, which is probably more valuable to them.
And I believe that this framework still has some room for optimization for generation with high batch sizes.
xyzzy_plugh 9 hours ago [-]
Your math is a bit funny if you're assuming the 1/0.5/98.5 ratios: you doubled input and output tokens but not cached. If you double cached tokens to match your original ratio it works out to around $1.11, and if you 10x the cached token cost it's around $6.08.
Based on your $0.83 estimate, the margin isn't great. This is within shooting distance of "at cost" which is probably pretty close to what DeepSeek is operating with, ignoring the value of the data they're collecting of course.
> And I believe that this framework still has some room for optimization for generation with high batch sizes.
If that optimization can bring this scenario closer to $0.50 then it gets pretty compelling, otherwise I'm not confident.
gpugreg 7 hours ago [-]
Oh, I messed up. Half-way through, I thought it would be a good idea to double the numbers so I don't have to deal with half millions, but forgot to also double the 98.5. Unfortunately, I can not edit it anymore.
I think the margins of DeepSeek may be a bit better than with this vibe-coded framework here, since they had the liberty of optimizing their models for their own hardware.
At the time, open frameworks were not anywhere close to achieving that number. Not sure whether they caught up. The software wizards at DeepSeek are quite skilled.
throw10920 8 hours ago [-]
> should not discount that DeepSeek also gets paid in data, which is probably more valuable to them
That's agentic feedback loops for training, right? Any more detail on this, such as how they actually tell whether that data is good or not? That seems like a very hard problem, and like the value of that data is low compared to just building their own, controlled RL gyms.
gpugreg 7 hours ago [-]
Agents usually start with ingesting the existing code base, and DeepSeek can use those code bases for pretraining. And they will have filters on top of that to throw out garbage.
I am not sure how they are using the data for post-training, but there probably are ways to get signal out of it, e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with the generated code, or started a new session with the same starting point as before, i.e. they git-stashed.
Generally, you can train on data that is quite bad (e.g. the entire internet). It will still work, but take much longer compared to clean data.
Aurornis 7 hours ago [-]
> They are selling at a loss?
Definitely not. Inference is not as expensive to operate as many people seem to assume. The frontier labs are probably making a lot of money from selling tokens. It’s covering all of the R&D costs like salaries, collecting training material, and running the large training operations that costs a lot of money.
cheema33 4 hours ago [-]
> It’s covering all of the R&D costs like salaries, collecting training material, and running the large training operations that costs a lot of money.
Are you claiming that the frontier labs like OpenAI and Anthropic are actually making a profit contrary to all the claims?
qwytw 2 hours ago [-]
They claimed that OpenAI and Anthropic have positive gross margins. I don't think there are many credible claims saying that's not the case (at least for API usage)?
dietr1ch 10 hours ago [-]
They claim their advantage is knowing how to serve their models efficiently, which is quite possible since they design for it.
pama 10 hours ago [-]
Use nvidia hardware instead and use a larger cluster serving many more users concurrently. Easily 10x–20x higher token rate per GPU with public solutions like dynamo and sglang.
simlevesque 8 hours ago [-]
They get all our invaluable data which they'll use to train the next model, to get more data, to train the model after.
jingpostmedia 7 hours ago [-]
[flagged]
drob518 8 hours ago [-]
I think Deepseek is selling roughly at cost (perhaps a slight premium). They don’t guarantee that they don’t train on the submitted prompts, so I suspect they are mining the data. Mining for what? Well, who knows. Best case, mining to make Deepseek better. That said, I use Deepseek all the time. It has done a whole lot of ‘ls’ commands on my system, though.
thrownaway561 10 hours ago [-]
This is exactly what I came to say. The price of Flash is so cheap that trying to run it locally or with your own hardware is pointless. I was using it about a month ago to program some stuff and ran it for 4 days non-stop and it cost me about $2.
ak_t 2 hours ago [-]
It's expensive to run it locally at full quality, but at least on my setup, its about 5 times faster than any API, and is completely private.
NitpickLawyer 9 hours ago [-]
> trying to run it locally or with your own hardware is pointless.
Serving local models has advantages other than price. If you work in restricted industries, or have a strong need to protect your IP, or if you just value privacy more than cost, you now have options.
YKutsev 4 hours ago [-]
[flagged]
ux266478 9 hours ago [-]
If you don't do any attention steering, custom decoding or meddle with the weights maybe. Services are worthless unless all you do is write positive prompts.
As others have mentioned, there's the privacy factor as well.
jorvi 9 hours ago [-]
With the cost of electricity, hardware depreciation and tok/s it rarely makes sense to run locally.
Tepix 9 hours ago [-]
If you have 2x DGX Spark it will run quite nicely. They cost only $8000 or so and use less power so you may be able to rent them cheaper than the MI300X.
The MI300X will vastly outperform it for only a slightly higher price.
K0IN 6 hours ago [-]
As other calculated, even with a Mi300 you could not saturae it enough with one stream to break even with the DS API, so I think renting sparks would make it even harder cause they are considerably slower.
(274gb/s vs 5.3tb/s)
langs 11 hours ago [-]
You need to optimize the KVCache part(save to disk to save compute) to achieve this goal.
Lwerewolf 11 hours ago [-]
The MI350p exists and should run a decent quant (say, the ~96GB antirez mix) well, but you can get two rtx pro 6000s for one of these, or 8x (actually more) r9700 + probably the gear to run them, etc.
Otherwise, you can probably buy one of these second hand from somewhere (SXM A100s are available that way) and run it in an adapter board.
touisteur 8 hours ago [-]
I thought MI350P wasn't available yet, curious where to source it right now.
FuriouslyAdrift 4 hours ago [-]
There's at least one systems integrator selling a rack server with 2x MI350Ps.
I haven't seen the cards all by themselves yet.
throwawayffffas 8 hours ago [-]
You can get one on ebay for like 20k, but it comes without the backplane and i dont think there is a pcie card adaptor from china like the ones for h200.
_joel 10 hours ago [-]
I thought it was a consumer grade GPU until I saw the 192GB of HBM and 256GB or RAM.
varispeed 10 hours ago [-]
To be fair the development of GPUs have stalled over the years. If they kept up with the progress instead of focusing on enterprise market, likely 256GB consumer GPU would be a norm today.
FuriouslyAdrift 4 hours ago [-]
It's a chopped down MI350X (roughly half the performance)
FuriouslyAdrift 1 hours ago [-]
Basically right from Lisa Su's speech: "AMD is essentially taking one of its MI350X accelerators and cutting it in half, resulting in a card with half as many compute resources, half as much memory, and perhaps most importantly, a bit over half of the power consumption"
baalimago 11 hours ago [-]
Give it an AI-bubble pop and these will be flooding the market.
segmondy 9 hours ago [-]
no they won't , the bubble is a financial thing.
the demand is real and not going away.
atwrk 7 hours ago [-]
The big question is whether the demand will stay if the subsidized pricing ends. That's what the bubble talk is about. Right now all the players compete for market share and don't care about the losses (hence the debt). But what happens if no one wants to lend them anymore?
jack_pp 6 hours ago [-]
I don't think inference is subsidized, it's the training. So what happens is, there's no new models anymore or are released slower.
FeepingCreature 4 hours ago [-]
API inference is probably not subsidized. Coding plans absolutely are.
vehemenz 6 hours ago [-]
Good distinction. The demand is partially driven by the low costs, which are only low because the major providers are losing money.
qwytw 2 hours ago [-]
There is no evidence they are losing money on inference, though?
Also if they are keeping the price low because they want to gain market share and reduce the competitiveness of Chinese models they won't be able to raise prices without providers serving open models (at cost + low margin) severely undercutting them.
_factor 11 hours ago [-]
They will be instantly bought out by companies, not individuals. The consumer bubble won’t pop for quite a while yet. Production also won’t ramp up while lack of real competition keeps the demand high.
aurareturn 11 hours ago [-]
When is it popping? Is the AI bubble in the room with us now?
dghlsakjg 8 hours ago [-]
Nvidia has ever so slightly underperformed the SP500 YTD (at the exact time this comment is being typed), so its basically the apocalypse already.
amrit3128 10 hours ago [-]
Tomorrow? Next year? In 5 years? Nobody can say. But we do know that AI is overvalued, so it WILL pop.
aurareturn 10 hours ago [-]
Well, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued?
I can tell you it will pop in 10 years and when it pops, it will still be 20x bigger than in 2026. Does that even make any sense?
People said AI bubble will pop soon in 2024 and that it was overvalued. Turns out, many AI stocks 10x, 20x since 2024. Actual usage has gone exponential as well. Anthropic revenue went from $100m ARR at start of 2024 to $80b ARR today.
ekidd 9 hours ago [-]
> Well, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued?
Well, given the literal trillions being spent, the only ways this pays off are:
1. AI replaces a non-trivial fraction of human employees.
2. Someone builds a Culture Mind, and humans become (hopefully) pampered pets of AIs we don't understand. Seems unlikely, but it would arguably count as a payoff even if it made money meaningless.
Or maybe the AIs don't want pets, and you get SkyNet. Which definitely doesn't care about paying off anyone's investments.
When you look at various news articles about investors, yeah, there are definitely a lot of rich people who think that they're going to automate all human labor or just bring about the Singularity. Possibly with them in charge of the rest of us. If you don't make these kinds of wild assumptions, then yeah, this is looking like one of the biggest bubbles ever.
aurareturn 9 hours ago [-]
Can we see some actual numbers, projections, models instead of vibes?
Interest alone, at assumed 5%, amounts to about $150 billions per year. That's probably higher than the combined AI revenue of the top 3 providers.
qwytw 2 hours ago [-]
That's what you should ask the companies spending massive amounts of money on AI though?
CamperBob2 7 hours ago [-]
1. AI replaces a non-trivial fraction of human employees.
I mean, it's pretty clear that's going to happen. How could it not?
The only question is whether resources are optimally allocated at the moment to prepare for this. That seems unlikely at best. So yes, there is probably a bubble, and if so, then yes, it will pop, and then life will go on, with resources better allocated. Just like when the dot-com bubble popped.
Joel_Mckay 9 hours ago [-]
Many are saying July 2027, as in the past these market corrections have correlated with Shrek movie releases.
Debt-backed investors have to pay up eventually. =3
baalimago 10 hours ago [-]
Next month perpetually
slaw 10 hours ago [-]
The AI bubble will pop when China gets access to EUV, so the earliest it could happen is 2030
tamimio 9 hours ago [-]
Thing is, GPUs will always be on demand, look at their history, initially for gaming, then for hash cracking, then 3D rendering, then for crypto mining, and now AI training and fine tuning. When AI bubble bursts, there will be another bubble taking over.
The only solution is more companies making high end units, only competition will make it better for consumers.
fergusfinn 8 hours ago [-]
nice! i think the higher HBM on Mi300x is really useful for this kind of thing
Strange that in the prior art they didn't list DwarfStar, as it is able to run the same model (probably quantized differently though) in less memory. Maybe the author isn't aware of it?
MaKey 7 hours ago [-]
The focus of this repo is the MI300X and DwarfStar doesn't include any optimizations/fixes for it.
Unfortunately, the MI300X is an OAM module. The MI350P is the one you want: It's a PCIe card, but it has less memory: 144GB.
Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized.
WhitneyLand 8 hours ago [-]
How do you figure that?
When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB.
And this implementation is already cutting down the 1M token context window you would normally get.
Tepix 3 hours ago [-]
For sure if you want to properly utilize the model with several users in parallel and large context you'll want two MI350P.
craftkiller 5 hours ago [-]
Just want to add that while the MI350P is a PCIe card, it is designed for servers. It just has a heatsink (with no fan) which the powerful full-case fans of a rackmount server are supposed to cool. So while the MI350P is certainly more attainable for us regular folk due to its formfactor, we won't be able to just drop it into our gaming PCs like a regular graphics card.
That being said, if you're dropping tens of thousands of dollars on graphics cards then picking up a rackmount case to go around the card is pretty insignificant.
cyberax 4 hours ago [-]
Just put a fan on it. It's just 600W, so nothing super-special is needed. Or add a water cooler.
monster_truck 4 hours ago [-]
You're going to need at least a couple loud ass IPPC 3000s to usefully move that kind of heat if you don't want it to throttle. And then another normal sized fan for the doorway of the room it's in.
Not exactly super special, but a ~constant 600W+ of heat tends to be a learning experience. It's worse than a high end gaming rig, much closer to a literal space heater. I don't work during the summer because it sucks fighting both this and the sun with AC. I do have a fan that slots into the window and can push or pull, but kicking the waste heat outside doesn't help when its humid.
WhitneyLand 8 hours ago [-]
Another headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”.
Dumbed down quantization?
No. Full intended inference weights preserved, so far so good.
Slow performance?
No again. Looks like you could get over 150 tokens/second.
Give up context window size?
Yes. Original model is trained for and served at 1M, this is 256k. A very practical tradeoff though. Codex is in this range, and quality does start to drop off toward the full size.
monster_truck 4 hours ago [-]
In my experience the 1M context is genuinely too much. The first time I swapped from OAI to DSv4P, I checked and double checked that the harness/etc was working correctly over the course of hours and hours of work thinking that I had set something up wrong because it simply never had to compact! The drop in quality is arguably less than that of what you get from compact to compact on Codex, which is good for what it is or was.
Was also surprised to learn just how much of Codex's window was being burnt on shit I didn't want or use. Sure I can pass this and that flag to eliminate most of it, but for a $200/mo product aimed at professionals, that isn't something anyone should have to janitor (also totally ignoring the bandaid of banked resets they've slapped over their repeated mistakes).
It's wild just how far $20 will get you with Deepseek, even at their new rates. Buyers Remorse is my very least favorite feeling, I felt sick thinking about what the $1200 I had given OAI this year would have gotten me had I only tried sooner.
bwfan123 7 hours ago [-]
I am curious if there has been work to remove experts from an open-weights model. The goal would be to reduce the size to be able to run on desktop GPUs without compromising quality. For a focused usecase - say coding, you dont need a model that knows world history. And, I am not talking about quantization. If it is possible to determine which experts are active for some usecases, and surgically remove the others.
smallerize 5 hours ago [-]
Experts aren't trained on separate tasks. More recent routers are designed to spread out requests even more evenly, and they were already pretty even.
Tepix 3 hours ago [-]
Yes. It's called REAP and from what I've seen, results aren't stellar.
monster_truck 4 hours ago [-]
Glazing over a lot, that's how they work already, just not in the way you think. A relatively small fraction of the model is active at any given time
xorfish 9 hours ago [-]
This is still quite a bit away from the performance that deepseek gets on their H800. In their DSpark paper they report a throughput of 15k tokens/s/gpu. The MI300 should be able to compete with the H800 so there are probably still quite a few optimizations that can be made.
somnial 6 hours ago [-]
throughput scales superlinearly with number of GPUs when networked well and deployed with wideEP, so 1x won't compare.
also it would be interesting to figure from the DSpark paper whether their numbers are consistent with the GPUs still being H800s, since they never actually say...
sylware 9 hours ago [-]
Is their hardware programming interface reasonable for implementing inference of frontier models: no quantization, several tera params?
BTW, how many many params open weight frontier models have? A few teras, 100s of teras?
wmf 4 hours ago [-]
Yes, ROCm can be used to run frontier models and is being used by OpenAI, Anthropic, and Meta.
sylware 2 hours ago [-]
I would prefer direct hardware kernel interface.
Like linux DMABUFs with userland hardware command ring buffers (I guess this hardware ring buffer instance would be specific to a VMID and a PASID).
wmf 2 hours ago [-]
You can do that (tinygrad style) if you want.
wren6991 8 hours ago [-]
Kimi-K3: 2.8T
Qwen3.8-Max: 2.4T
DeepSeek V4 Pro: 1.6T
DeepSeek V4 Flash: 284B
(all are total parameter counts, not active parameters)
sylware 3 hours ago [-]
Rumors say chatgpt/claude/gemini/etc are in the 100s of teras. True?
wren6991 2 hours ago [-]
I'll ask my uncle (he works for Nintendo) and get back to you on that one
sylware 2 hours ago [-]
My question is that wrong?
wren6991 2 hours ago [-]
Sorry if the joke didn't land; I have heard a lot of different numbers for the size of US labs' models, but never seen any of them substantiated, so I think you're likely to just get more rumours in answer to this question.
My personal take, with no sources: 100T sounds excessively high given they need to be able to actually serve these things on commercially available hardware. I would guess they are in the same order of magnitude as the Chinese frontier models. It's possible their edge is in RL training methods, training-time compute, and access to data (e.g. from customers' CC/Codex sessions), not in model size.
Edit: Now I think about it, this might be the cheapest way to run the DeepSeek V4 Flash 0731 on a dedicated inference server at original weights. I haven’t run mixed load benchmarks but I guess it’s possible to generate $3-$4 worth of tokens per hour and still maintain a usable per-user throughput.
To be clear this project doesn’t aim to achieve the best inference economics per token. MI300X doesn’t have native MXFP4 so it’s not even the right platform for the model. That’s why very few deployment recipes are available.
It’s interesting to me because MI300X is quite accessible to a small team with budget for just 1-2 GPUs. DeepSeek V4 Flash otherwise wouldn’t even fit on 2x H100s.
We can run several coding agents during the day and batch inference jobs overnight and serve the entire team with guaranteed privacy, without compromising precision or speed.
In fact we found that many inference providers are quantising the weights or even KV cache, and due to the low prices they serve at massive batches, resulting in unstable throughput. I ran GSM8K as a quick validation test and this deployment is “better” than the OpenRouter endpoint in a statistically significant way (I wouldn’t name the provider here). I will run some follow up benchmarks and update the repo when I find some time.
You shouldn't rent one out if you're just serving it for yourself, but from a financial standpoint if you sell to users you can take a 100% margin.
This is not the real margins, even if you are selling to 8 users it's 90 tps per median stream. So assuming that .6-.7$
This is not even remotely worth it.
You need to 3x this tps(~1500 tps) to be worth it, and that's what most providers are doing, at 20-30 users at 50-60 tps with better optimized batch processing and kernels you can make some profit.
Lets assume that you've got 2 million input tokens, 1 million output tokens and 98.5 million cached tokens to process. That would cost 2 * $0.14 + 1 * $0.28 + 98.5 * $0.0028 = $0.8358 with DeepSeek API pricing.
For comparison, it would take 2M / 8000 + 1M / 800 = 1500 seconds to process this amount of tokens with the linked framework, which is about $0.83 when we assume $2/hr for one MI300X.
However, other inference providers have 10 times higher prices for cached tokens, which results in a comfortable margin.
And we should not discount that DeepSeek also gets paid in data, which is probably more valuable to them.
And I believe that this framework still has some room for optimization for generation with high batch sizes.
Based on your $0.83 estimate, the margin isn't great. This is within shooting distance of "at cost" which is probably pretty close to what DeepSeek is operating with, ignoring the value of the data they're collecting of course.
> And I believe that this framework still has some room for optimization for generation with high batch sizes.
If that optimization can bring this scenario closer to $0.50 then it gets pretty compelling, otherwise I'm not confident.
I think the margins of DeepSeek may be a bit better than with this vibe-coded framework here, since they had the liberty of optimizing their models for their own hardware.
For DeepSeek V3, they claimed a cost profit margin of 545%: https://github.com/deepseek-ai/open-infra-index/blob/main/20...
At the time, open frameworks were not anywhere close to achieving that number. Not sure whether they caught up. The software wizards at DeepSeek are quite skilled.
That's agentic feedback loops for training, right? Any more detail on this, such as how they actually tell whether that data is good or not? That seems like a very hard problem, and like the value of that data is low compared to just building their own, controlled RL gyms.
I am not sure how they are using the data for post-training, but there probably are ways to get signal out of it, e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with the generated code, or started a new session with the same starting point as before, i.e. they git-stashed.
Generally, you can train on data that is quite bad (e.g. the entire internet). It will still work, but take much longer compared to clean data.
Definitely not. Inference is not as expensive to operate as many people seem to assume. The frontier labs are probably making a lot of money from selling tokens. It’s covering all of the R&D costs like salaries, collecting training material, and running the large training operations that costs a lot of money.
Are you claiming that the frontier labs like OpenAI and Anthropic are actually making a profit contrary to all the claims?
Serving local models has advantages other than price. If you work in restricted industries, or have a strong need to protect your IP, or if you just value privacy more than cost, you now have options.
As others have mentioned, there's the privacy factor as well.
I found an offer to rent two at $1.65 per hour https://spark.enverge.ai/#pricing
The MI300X will vastly outperform it for only a slightly higher price.
(274gb/s vs 5.3tb/s)
Otherwise, you can probably buy one of these second hand from somewhere (SXM A100s are available that way) and run it in an adapter board.
I haven't seen the cards all by themselves yet.
Also if they are keeping the price low because they want to gain market share and reduce the competitiveness of Chinese models they won't be able to raise prices without providers serving open models (at cost + low margin) severely undercutting them.
I can tell you it will pop in 10 years and when it pops, it will still be 20x bigger than in 2026. Does that even make any sense?
People said AI bubble will pop soon in 2024 and that it was overvalued. Turns out, many AI stocks 10x, 20x since 2024. Actual usage has gone exponential as well. Anthropic revenue went from $100m ARR at start of 2024 to $80b ARR today.
Well, given the literal trillions being spent, the only ways this pays off are:
1. AI replaces a non-trivial fraction of human employees.
2. Someone builds a Culture Mind, and humans become (hopefully) pampered pets of AIs we don't understand. Seems unlikely, but it would arguably count as a payoff even if it made money meaningless.
Or maybe the AIs don't want pets, and you get SkyNet. Which definitely doesn't care about paying off anyone's investments.
When you look at various news articles about investors, yeah, there are definitely a lot of rich people who think that they're going to automate all human labor or just bring about the Singularity. Possibly with them in charge of the rest of us. If you don't make these kinds of wild assumptions, then yeah, this is looking like one of the biggest bubbles ever.
Interest alone, at assumed 5%, amounts to about $150 billions per year. That's probably higher than the combined AI revenue of the top 3 providers.
I mean, it's pretty clear that's going to happen. How could it not?
The only question is whether resources are optimally allocated at the moment to prepare for this. That seems unlikely at best. So yes, there is probably a bubble, and if so, then yes, it will pop, and then life will go on, with resources better allocated. Just like when the dot-com bubble popped.
Debt-backed investors have to pay up eventually. =3
The only solution is more companies making high end units, only competition will make it better for consumers.
we did some work on this for 2xMi300x (kindly referenced in the readme) https://blog.doubleword.ai/deepseek-v4-flash-mi300x. https://hotaisle.xyz/quick-start hotaisle is great for getting Mi300x to experiment with
I assume this is parallel work.
Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized.
When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB.
And this implementation is already cutting down the 1M token context window you would normally get.
That being said, if you're dropping tens of thousands of dollars on graphics cards then picking up a rackmount case to go around the card is pretty insignificant.
Not exactly super special, but a ~constant 600W+ of heat tends to be a learning experience. It's worse than a high end gaming rig, much closer to a literal space heater. I don't work during the summer because it sucks fighting both this and the sun with AC. I do have a fan that slots into the window and can push or pull, but kicking the waste heat outside doesn't help when its humid.
Dumbed down quantization?
No. Full intended inference weights preserved, so far so good.
Slow performance?
No again. Looks like you could get over 150 tokens/second.
Give up context window size?
Yes. Original model is trained for and served at 1M, this is 256k. A very practical tradeoff though. Codex is in this range, and quality does start to drop off toward the full size.
Was also surprised to learn just how much of Codex's window was being burnt on shit I didn't want or use. Sure I can pass this and that flag to eliminate most of it, but for a $200/mo product aimed at professionals, that isn't something anyone should have to janitor (also totally ignoring the bandaid of banked resets they've slapped over their repeated mistakes).
It's wild just how far $20 will get you with Deepseek, even at their new rates. Buyers Remorse is my very least favorite feeling, I felt sick thinking about what the $1200 I had given OAI this year would have gotten me had I only tried sooner.
also it would be interesting to figure from the DSpark paper whether their numbers are consistent with the GPUs still being H800s, since they never actually say...
BTW, how many many params open weight frontier models have? A few teras, 100s of teras?
Like linux DMABUFs with userland hardware command ring buffers (I guess this hardware ring buffer instance would be specific to a VMID and a PASID).
Qwen3.8-Max: 2.4T
DeepSeek V4 Pro: 1.6T
DeepSeek V4 Flash: 284B
(all are total parameter counts, not active parameters)
My personal take, with no sources: 100T sounds excessively high given they need to be able to actually serve these things on commercially available hardware. I would guess they are in the same order of magnitude as the Chinese frontier models. It's possible their edge is in RL training methods, training-time compute, and access to data (e.g. from customers' CC/Codex sessions), not in model size.