Once it stops being subsidized, and prices go up 30-50%, there will definitely be a correction. It's also a race to the bottom - frontier models will continue to improve but "good enough" has already been hit, and at that point companies will need to compete on price, and they can't go too much lower in pricing without losing money.
That said they ARE building new fabs, that shit just takes years.
It’s also largely good enough now that you can do coding (which is a big driver of usage) on a local model. It’s slower, but so what. Give it a prompt and go make a cup of coffee, vs give it a prompt and instantly get your code. Still faster than doing it yourself.
This guy codes, and as soon as I can run a decent model with 8 gpu (the max a commercial motherboard can take), I am 100% getting some uncensored model and using it for my main LLM, MAYBE with the occasion claude request ($20 per month) for ultra-complex stuff). Bye bye chatgpt and gemini, I wont miss you.
I haven’t played with open router yet but apparently that’s the killer feature: use a local model most of the time, automatically switch to a paid one when needed.
Once it stops being subsidized, and prices go up 30-50%, there will definitely be a correction. It's also a race to the bottom - frontier models will continue to improve but "good enough" has already been hit, and at that point companies will need to compete on price, and they can't go too much lower in pricing without losing money.
That said they ARE building new fabs, that shit just takes years.
I can tell you for a fact that the $200 ChatGPT plan can end up costing them $17,000............................
FUCK GAMING, STOCKPILE AMMO!
Are you lost? You don't belong here.
It’s also largely good enough now that you can do coding (which is a big driver of usage) on a local model. It’s slower, but so what. Give it a prompt and go make a cup of coffee, vs give it a prompt and instantly get your code. Still faster than doing it yourself.
This guy codes, and as soon as I can run a decent model with 8 gpu (the max a commercial motherboard can take), I am 100% getting some uncensored model and using it for my main LLM, MAYBE with the occasion claude request ($20 per month) for ultra-complex stuff). Bye bye chatgpt and gemini, I wont miss you.
I haven’t played with open router yet but apparently that’s the killer feature: use a local model most of the time, automatically switch to a paid one when needed.
ive not checked in in a while, what are the best local models right now?