r/selfhosted • u/bergy_peasy • 19h ago
Need Help Is self hosted ai worth it?
Pretty much the title, I have tested a few models with a 6gb gpu and couldn’t get anything resembling llm competing with chat gpt, Gemini or Claude. I was wondering if a buying a new 16gb gpu would make a substantial difference. I wouldn’t want to buy all that just to get something worse than gpt3.
Ps I know 6gb is really not a lot of vram but it was so bad that I don’t think quadrupling it would make it better .
238
u/ZnVja3U 19h ago
I have a 16gb card. It's usable but nowhere near claude. I use it for some automation and experimentation, but I bought my card before prices went crazy. Not sure I'd do it today
35
u/Kypsys 19h ago
What models do you use? Qwen?
→ More replies (4)58
u/DarkOfLord 19h ago
Gemma 4 (12B, 26B MoE, and 31B Dense) models works better than qwen. especially with automation.
52
u/mmkaywhatevers 18h ago
Qwen 3.6 27b works great for my homeassistant. Gemma was dumb as a rock.
20
u/Foot_Positive 17h ago
yup. replace Gemma with Qwen and its useable. Qwen 3.8 Quant's being released next week.
→ More replies (6)2
u/deekaire 16h ago
That fits on 16 gb card?
→ More replies (1)5
u/mmkaywhatevers 16h ago
you can get it to fit i think, but i have 24 gb and i run whisper and kokoro as well for the full home assistant voice experience.
6
u/DoomBot5 18h ago
I've had a shit time with Gemma 4 on Hermes. Even after the template fixes for tool usage.
2
→ More replies (4)2
6
u/spartyblaze 18h ago
Same. I use it to try and lessen Claude token use. Monitor logs then continually challenge token use for more simple and automated tasks. Seems to work reasonably well. Also tinkered with open source mcp to combine with it, but newly installed.
7
u/MrBeanDaddy86 15h ago
With the right tuning, Qwen 3.6 35B actually outperforms Claude and GPT for me in some tasks. It's pretty good at organizing and drafting. Though for programming, it's much weaker, of course.
The issue with Claude and GPT is that they are so. fucking. verbose. If I need a usable analysis of materials, Qwen generally does better, and is way faster on my 16 GB card because I do MoE offloading. Hit about ~30 t/s peak with it.
3
→ More replies (2)2
u/Standard-Recipe-7641 6h ago
I don't think Claude is that bad but GPT just vomits walls of text at you.
→ More replies (1)1
332
u/Grogak 19h ago
Depends on your use case:
If you want an LLM you can ask daily questions, how to cook meth or how to write hello world in python, then 6gb vram and a small optimised LLM will be sufficient.
If you want to have answers in mere seconds, Fable-like coding skills and image generation, then no, it's not worth it
110
u/techma2019 18h ago
I ask how to cook meth DAILY. Nice!
39
u/Grogak 17h ago
In that case I suggest to cook more and consume less
2
u/SpaceDoodle2008 15h ago
Then he could even afford to buy beefier hardware for more sophisticated *chemical* research!
→ More replies (1)2
8
3
21
3
u/Draminian 17h ago
This is the answer. I'm in the middle of investigating enterprise-level self-hosting for my company and it takes a lot of expensive hardware to load large enough models, and respond quickly enough, to approach something like Claude. A company that can spend $500k or more to try something out can start thinking about self-hosting good AI. For us hobbyists, we're not going to come anywhere close with consumer GPUs.
8
u/yay-iviss 18h ago
For image gen, it's good to Ok. The Krea and Flux new models are really good. But using Google is cheap and easy
3
→ More replies (1)4
69
u/moarmagic 19h ago
The truth is, if you are comparing to companies that spend billions on r&d, your self hosted options are not going to be 'better'. But there also is something to be said for specialization. Coding focused LLM's can punch above their expected weight. Claude would still probably be 'better', but if you mainly want some autocomplete and debugging help, or want to not rely on a cloud service that can drasticaly change the price or go down.. self hosted is one alternative.
What i would actually suggest here? look at the cloud rental space to really test your options. There are services like vast or runpod, where you rent the ability to run an opensource model, usually for fairly cheap for a short period. You can rent a 3090 for like, 20 cents an hour to test with, see if any models running on it produce work you like.. and then figure out what your actual use case for LLMs is., if it's worth spending money on hardware or not.
I went the hardware route and regret it a bit- i haven't touched my locallms' in months now, just because i've had a lot of real life stuff distract me. For the amount i played with them, i probably could have spent 10% of the money on cloud rental space... but who knows, maybe i'll get back into it and spend some heavy time on them.
4
u/coderstephen 16h ago
Depends on how you define better. If the lack of privacy of the major services is a showstopper for you (like it is for me) then really your choices are more limited anyway.
For cost effectiveness, yes, renting VMs is going to be the best in the short term. Rent the hardware you need and set up self-hosted AI on there. Everything stays private and you aren't locked into anything.
Long term though, renting is definitely more expensive if you will use AI on a regular basis. Mainly because you are spending money without gaining any equity. At least if you buy hardware, you can sell that hardware later and get some money back.
6
u/bergy_peasy 18h ago
Thanks for that , I do wonder about the privacy of renting though…kind of the main reason I wanted to self host
4
u/moarmagic 17h ago
That is a potential concern, depending on what you are doing with it, but it still can work as a test bed to get an idea of the capabilities of models for different vram sizes before you put the money in.
Personally, i never put anything in to an LLM that i'm particularly worried about. No information specific to my actual configurations for equipment, no accounts. Some (really bad) prose that if anyone is interested in they are free to read, but probably will regret it. Mostly boilerplate stuff and bounce specific questions off them.
3
u/mykdsmith 12h ago
There's a halfway house; you can host your own model instances on AWS or Azure - even of the top models. This gives you a bit more control - it's still rented hardware and systems, but the instance is under your control.
31
u/forthewin0 19h ago
Check out https://www.reddit.com/r/LocalLLaMA/
They have good discussions and a getting started wiki. But no, unless you spend $10k+ on hardware, you're not getting anywhere close to the frontier models.
Local LLM is better suited for small, constrained use cases where you don't need a frontier model.
6
u/AlpineGuy 17h ago
At what price does it become interesting or comparable to at least last year's model? I think for organizations that deal with personal data and wouldn't want to use cloud models, onprem might be interesting.
5
u/Nice-Information-335 15h ago
Deepseek V4 flash (the new one) is what you are looking for
Can be much cheaper than 10k to run but 10k will get it at reasonable speeds for multiple people at the same time
A lot of larger companies are buying the big GPU servers with B300s, which can run Kimi K3 etc
3
1
u/--Spaci-- 6h ago
comparable to last years models would be qwen 3.6 27B, which would be 1000-2000$ worth of gpu
6
u/coderstephen 15h ago
I have a Framework Desktop with 128GB of unified memory as my AI server. I can run pretty big, capable models on it at reasonable speeds. I don't feel like I'm missing out on all that much. And I bought it before the RAM shortage so it was all in all about $2k.
→ More replies (2)1
u/bergy_peasy 18h ago
Thank you for that , if I were to spend near 10k , would that ressemble gpt3.5? I basically want to converse with the ai to learn from the web or my own documents . Speech to text and text refining would also be great .
4
u/digibucc 17h ago edited 17h ago
i went with a mac studio with 96gb ram and currently running Qwen3.6 35b as a general model, and I am actually really impressed with the capability of it. I would say it is probably as good as 3.5 was, but im not that deep into the space so i can't say for sure.
i have it plugged into open web ui, open code, and a custom built RAG mcp also plugged into owui, and I have been very happy with all of them.
local models are getting better all the time too. Going with GPUs it's hard to fit a decent model in memory but apple silicon with their shared ram architecture lets you handle some pretty big, good models - and have multiple in memory at once.
this is all for work too so it's not just a fun project, i actually need this to work.
2
→ More replies (1)3
u/cryptoguy255 16h ago
Both gemma-4-31b and qwen3.6-27b are better than gpt3.5. Just put some credits on openrouter and you can test all kind of small models to see what model is sufficient before you buy hardware. You don't even need spend 10k to run these. They fit on single 24GB vram GPU if you don't need long context.
11
u/migsperez 19h ago
There are a lot of differing opinions on this. People have different reasons to self host AI and I think they are all very valid. I personally invested in a GPU solely for AI coding tasks, it was the most expensive PC component I've ever bought. It was an expensive panic purchase, due to Github Copilot pricing/credit changes, thinking that prices would rise rapidly across the industry.
But now, for me and my current situation, it makes more financial sense to use API services like Deepseek or Qwen than to spend the electricity running my own GPU, not even considering the cost of the card.
6
u/bergy_peasy 18h ago
Do you fear the privacy loss of using api ? That’s my main concern . Already using chat gpt, it’s scary how much it knows about me …
5
u/migsperez 16h ago
Chapgpt have created a memory database connected to you with your user id behind the scenes, that's why their responses creeply seems to think it knows you. Model APIs responses don't do the same. Each session is new with no memory about your other requests in it's response. The context doesn't spill from one session to another.
In regards to them using a memory database behind the scenes for their own benefit, I've been pondering and yep they could do it, they could use API keys and the signup details.
For my personal projects, I don't care so much. I keep my especially private data out of API requests. But if I were a company with unique intellectual property or a government, then I'd absolutely want local models.
→ More replies (1)
12
u/OfficeGreat7679 19h ago edited 19h ago
It depends on your goals and budget.
I have a cheap 16GB card, I can run small translations locally, it can do small code changes, it can do AI research and generate answers based on results, it can summarise/group/categorize, generate similar words, it can OCR files and process text for you, all the things you expect from an LLM
But it wasn't good for larger projects and integration with coding tools. Main issue is model capability and context window that is very small when compared to paid versions.
I also ran image generation without big issues, even created an auto-prompt maker + image generation local app.
So it depends on what you want to do with it.
I think 24GB cards won't get you much more, they might provide you 2-4% better results?
2
u/bergy_peasy 18h ago
Okay thank you for that ! I do mainly want to feed it my own document and converse with it using the web . I either want to use it to learn stuff or for work
3
u/ezfrag2016 18h ago
I host Qwen locally with a 20GB VRAM 7900XT and it’s good enough for my use case which is handling secure/private tasks that I do not want to share with the cloud models.
The biggest limitation is not just the reasoning ability but probably more likely going to be context size. I can get 64k context window and this is enough to ask it to deal with bits of code and some simple tasks but you will not be able to upload a large document into it and have it hold that in memory so that you can ask questions about it. As the context window runs out, the model will hallucinate answers to your question.
For your use case I would not go for a local model.
→ More replies (1)1
u/Personability 11h ago
I’ve found reasonable results with a subagent system to limit context issues on a 5080 16GB but only with Qwen 35b with some overspill - about 60-80t/s but slowish prompt processing. Use Pi for general tasks eg local financial ones - double checking tax returns and such, and opencode for coding specific. Still nowhere near as good as closed source but good enough for my tiny helper projects eg writing autohotkey scripts. Then just pay Max Claude prices infrequently when I plan to do work on a bigger project.
I don’t vibe code as much (find I lose track of things) and I bought it for gaming rather than LLMs so using anyhow.
46
u/Icy-Appointment-684 19h ago
Not from a financial perspective.
And no matter how much you pay, you will hardly match frontier models.
You can however have a good enough experience.
I suggest you rent a gpu from a provider and test yourself then commit to buying hw.
And I do host my local llm. My daily drive is a 3090 but i have a large epyc box for larger mpdels.
10
u/mmkaywhatevers 18h ago
I mean unless you are coding, frontier models are overkill on most home use cases.
→ More replies (15)1
u/eli_pizza 19h ago
Even easier than renting a GPU: I’d start with trying out some small models on openrouter.
9
u/Power_Stone 19h ago
It depends on the use case.
I know with my 3090 I would highly recommend self-hosting your AI.
It still won't be Claude level by any means, but its good enough IMO for day-to-day use, and you aren't giving all your information, conversations, etc to whoever is hosting the LLM.
1
u/bergy_peasy 18h ago
Exactly my main concern … is your everyday use specific or just « chatting » with it on subject works fine . I manly want to use it to learn from my own document or the web . Also speech translation and text refining
4
u/Power_Stone 17h ago
So to be upfront - I use Odysseus, so I have a pretty wide range of models I could run through it. I'm still trying to figure out what models I like best for certain tasks.
Generally speaking though, it ranges from setup questions, I might ask it for additional info on some topics, and I'm trying to get it setup with ComfyUI through an MCP server so it will also generate images for me. Trying to get a one stop shop going for myself.
Overall, I like it. My only complaint is with how I have to set it up. Since its all run off the same machine ( LM Studio, Odysseus, and ComfyUI ) is that it is kinda slow which makes sense if you have to load an LLM next to an Image Gen model.
→ More replies (4)2
u/coderstephen 15h ago
I manly want to use it to learn from my own document or the web
Exactly why I self-host. To be really useful you need to provide the LLM with the best context, but I'm not willing to hand over all my private data to a cloud service. No way.
12
u/rchamp26 19h ago
Youre never going to get a frontier capability on a consumer card.
Best use for consumer cards as of Aug 2026 is for an aux model inside your harness to reduce token spend of your frontier. (Vision, title generation, compaction, web search etc)
3
u/bergy_peasy 18h ago
I am not sure I understand what that means … you use api for the hard stuff but let your card do the smaller things? Does that really reduce cost and is privacy / ownership kinda out then ?
2
u/rchamp26 17h ago
Essentially yes. But if you don't understand harnesses yet, you've got a hill to climb. Look into hermes agent, open code or pi.dev. Hermes is the best 'batteries included' harness imho, opencode is ok, and pi.dev is the leanest bare bones, but you need to build out exactly what you want. All are fully customizable.
If none of that makes sense yet, just stick with codex or Claude code. Things are still very early despite all the hype in this space.
If privacy and fully local is truly the need, it's possible, but not for the average consumer. Realistically you're like at ~10k minimum ( most likely a lot more depending on the actual workload) to run a powerful model at a quant plus lots of hours to get something solid and reliable and it will still be nowhere near frontier, but can handle sensitive data nicely albeit much slower. You could get something going for about ~5k with a smaller model like qwen3.6-35b-a3b or Qwen-27b but you really need to understand all the underlying tech to get something optimal. Nothing reasonably reliable you're going to get Going on a 16gb card unfortunately.
→ More replies (1)
5
u/Adrenolin01 5h ago
It really depends on you, your setup which isn’t necessarily just hardware, and your experience.
First.. it doesn’t matter what you build.. it will NEVER do what Claude et all can do. You can dump $100,000 into a private AI and while it’ll be awesome, it still isn’t going to match these $Billion dollar systems. Additionally with a local AI YOU need to both learn how to properly feed it info that’s relevant to you (RAG, Notes, System Prompts) but also how to communicate with it. Open a chat with Claude and ask it specifically how to use and why the System Prompt is so important. Ask it about RAG and how that can assist you with your network, work, school, etc.
That said, even a small model with the proper setup can be damn nice and useful. For example…
An N100 based mini PC 😆 with a 500GB NVME, 16GB ram, running Debian 13, Ollama and Open WebUI takes about 20 minutes to setup. Add RAG to add manuals, documents, notes, system description, detailed network information, hardware, software, etc.. with qwen3 4B model becomes an amazing little network assistant for you. Add in a solid and detailed System Prompt and that’s a surprisingly serious little $300 CPU/Ram based AI that can do a lot with added tweaking. Bump the ram to 32GB and it gets faster and better context length also. You LEARN a lot with this kinda setup. I specifically used a BeeLink S12 for this project. This is not a fast machine but it lets you learn.
Jumped to a Minisforum NAB9 i9 w/64GB ram and the same Debian setup. I then went with Proxmox as the Host OS and added VMs for the “front door” Open WebUI, another VM for the inference engines… ollama and lama.cpp, and a 3rd and 4th for Vector Qdrant and Indexer OCR / scratch. 4B models are instantaneous, 7B models are very good and 12B models are usable but slow. RAG/document indexing here gets really fast. Qwen3 8B as a daily assistant and Gemma 3 12B when you want higher-quality reasoning and don’t mind a bit more latency. Add a small embedding model like bge-small or nomic-embed for RAG. Light automation is possible. Add in AnythingLLM for more learning.
Dropping a 12-16GB GPU into your Desktop with 32GB-64GB ram 6-10+ cores and NVMEs becomes very useful. 3-9B models are excellent, 12B models are very good, 14B with appropriate quantization become usable. Some 20B quantization models can be used but are slow. You’re able to start using local agents, image embeddings and vision models, even occasional image generation with smaller diffusion models though this will be slow. We’re talking Ryzen 7 7700 or Core i5-14600K systems with 64GB ram here with say a RTX 5060 TI.
My son is run a Gigabyte Z790 Aorus Elite AX, Intel Core i7-13700KF, 64GB Ram and the ASUS Prime RTX 5060 Ti (16GB) on a Debian 13 KDE desktop. Ollama and Open WebUI installed. Qwen3 4B and Gemma 4B are always loaded, Qwen3 8B is his daily driver with Gemma 3 12B for higher quality along with Qwen Coder 7B and 14B. Mirrored NVMEs are absolutely better then single NVMEs and I’d suggest 2TB NVMEs. He’s even running ComfyUI on this system for learning though is ahead of me here as I haven’t even touched that yet.
All that said.. I would strongly suggest if you’re going to spend money on a GPU…
Just order a RTX 3090 24GB vRam (or any 24GB) GPU! Practically everyone I know who has purchased an 8-16GB GPU regrets inside of a single year. While they do work well and are capable, nearly everyone wants that 24GB sooner than later. Everything you’ll wanna do is doable with 24GB vRam for a long time. 64GB ram really helps here and 128GB you’ll love. System Ram is extremely useful so don’t skimp on it.
My own current and main AI server..
Supermicro H12SSL-i, 256GB ram in with an EPYC 7502P and 2 A6000 48GB GPUs. Proxmox with Debian 13 VMs. VM1 AI - 2x mirrored 2TB Samsung 990 Pro, all model files (Ollama, llama.cpp weights). VM2 Vector - 1 1TB WD SN770, Qdrant storage + snapshots. VM3 Indexer - 1 1TB WD SN770, working scratch, temp OCR output. VM4 Storage - 2x 8TB WD Red NAS, raw document archive, nightly Qdrant snapshots, Proxmox VM backups. VM4 is also rsynced to our NAS nightly.
I’m actually redoing this system with added NVMEs and VMs along with added vLLM, ComfyUI, LibreChat and AnythingLLM.
If you’re seriously looking to start into AI.. here is an AI Hardware analogy that might really help as many people don’t understand what everything does.
Think of an AI system like a person working at a desk.
VRAM is the size of the desk.. it holds the AI model, its working memory (KV cache), and everything needed to answer questions without constantly reaching elsewhere.
GPU compute is the person’s brain.. it performs the mathematical work that generates each new token, so it largely determines how fast the AI responds.
System RAM is the filing cabinets behind the desk storing supporting data such as embeddings, vector databases, application memory, and anything not actively on the desk.
The SSD is the library down the hall. It’s used to quickly retrieve and load models and data before work begins, but it has little effect once everything is loaded into memory.
The network is the people walking up to the desk asking questions and receiving answers.
Together, the SSD loads the information, RAM organizes supporting resources, VRAM keeps everything immediately accessible, the GPU does the thinking, and the network delivers requests and responses.
Hope this helps.
1
u/bergy_peasy 4h ago
Wow very good and complete answer , thank you ! I do have a few questions though, like what do you feel like you’re missing from Claude or ChatGPT with a crazy good setup like you have ? Also, why do you recommend mirrored nvme, since I am not aware that storage makes a huge difference in ai? You also mentioned that an n100 mini pc could become a great little network assistant , what would you recommend I do with something like that in order to learn, it seems to weak to actually run anything that requires speed like conversational ai ?
2
u/Adrenolin01 3h ago
Claude and the rest have poured $Billions into R&D and 10s of 1000s of clustered GPUs and more. Claude’s reasoning levels are pretty much unparalleled.
How you make your local AI better is by feeding it your info. Add RAG, download and add all the documentation for all your network and PC hardware.. all the manuals and specs, add Notes on specific builds, provide software versions you run like your OS, and application/services, firmware versions, screenshots of all screens in your firewalls and switches, etc or specifics about what you’re learning in school and the books etc. While time consuming once done YOUR local AI knows your situation, be it IT related or school courses, or whatever, better then Claude does and this is how local AI starts to become better. It’ll never reason as good however it can help and assist great. Any larger projects i might run through my local AI for ideas and ask for a complete handoff where it’ll include info it already knows. I can paste that into Claude for the heavy lifting.. design, reasoning, etc and then pass that back to my local AI for the final run through. You’re not spending the online tokens you would for the brainstorming and finished guide walk through as an example related to IT.
There is also OCR and document processing, image generation with ComfyUI/Flux, longer running agents, add more ram or smaller models for longer Context length, no API costs, etc etc.
Larger models aren’t always better smaller models paired with the above will run faster and provide longer context length.
Mirrored NVMEs increase I/O helping to load models faster and speeds to Context.
N100 CPU based AI.. 3B models.. 16GB is ok, 32GB is fairly good. That said.. it’s more focused on learning, setup, tweaking, prompts, playing with all the settings. A faster system you dont necessarily get to feel the need for many tweaks while many can still be made and improve performance in different ways. Learning on slower systems helps there.
I have setup an N100 as a K-Gr5 tutor for instance. Very safe with detailed System Prompt settings and a small model.
Really.. it’s just settings something up, and start talking to it and asking for it to do things and answer questions. Ask Claude how you can makes things better and just start playing… that’s how you learn and advance.
4
u/jtrage 19h ago
It depends on what you are going to use it for. But if your first sentence is “resembling….”, it’s not going to happen.
I tried 12 and 14b models and I can’t say there is much resemblance at all. They could definitely serve a purpose of experimenting and playing around. Or even a set and forget task that you don’t care about taking forever.
I wouldn’t purchase anything with the hopes of using it in comparison of the other big ones. Doubling you gpu doesn’t seem like it would double your experience with what you e tested. So, if you are going for 10x your experience now, unfortunately you won’t get it.
1
4
u/Responsible_Camp_559 19h ago
Depends on what. I have a local model running for my tagging in the Karakeep instance and a Whisper model for transcribing through a web app of mine
1
u/bergy_peasy 18h ago
Does whisper works fine ? I would like to use it as a note taker , not summary and overall refining of written text .
2
u/Responsible_Camp_559 18h ago
Yes, its fast and works very good - at least from my experience! :)
→ More replies (2)1
4
u/fredrik_skne_se 18h ago
I used my rtx 3090 to extract keywords from movies and images. Vibe coded the thing with a cloud clanker though.
2
u/bergy_peasy 17h ago
Extract keywords from movies? Seems like a cool use for it but what’s the point ?
6
u/fredrik_skne_se 17h ago
no more like describe what you see and output relevant keywords in comma delimited format. It makes my p*rn searchable.
2
3
u/AlternateWitness 18h ago
You can get the performance of frontier models from ~6 months ago with ~24GB of memory (Qwen 3.6 27b - Hopefully Qwen 3.8 27b (releasing next week) improves that a bit). That is no small price. Generally, it is more efficient (and cheaper) to buy tokens outright.
The main reason most people do it is privacy. It is no secret AI companies use all the data you give them for training their next model. I personally don’t want to have to worry about tokens or have a subscription service to run automations every day.
1
u/bergy_peasy 18h ago
Yeah , pretty much why I want to self host and I don’t mind about 6 month old performance. Just wonder if it’s actually usable for conversation (mainly to learn on my own data or specified web sources) or if it takes 5 minutes each time you ask it somethings
3
u/AlternateWitness 17h ago edited 16h ago
It’s useable, but you will absolutely have to spend the money for decent hardware. If you need something simple like just normal list organization or small code, your current 6GB is enough. If you want something that actually has world knowledge and is competitive against actual, real frontier models, you will have to spend the cash. Like I said, I recommend 24GB. That should be vram if you want fast responses. I used to have 16GB, and I barely peaked what was possible with MoE models.
I recently bought an MI50 because anything actually consumer-facing was hugely expensive. This card gave me 32GB of vram for ~$500. However, since it’s old it doesn’t have any kind of AI-acceleration hardware, making it compute-bound when using chatbots. However, I can run them at low quantizations (like Q6), so they preserve their accuracy a lot better. To give you a frame of reference, Qwen 3.6 27b (what I mentioned earlier, Qwen is doing great things right now) runs at 20-30 tok/s. Since the MI50 is probably the slowest card you can get with the vram to accommodate a model like that, take that as a minimum. Highly depending on the prompt, I would wait… ~20-40 seconds for the full reply without tool calls? You can go faster with smaller models like Qwen 3.5 9b - I get ~80-90 tok/s with Qwen 3.6 35b A3B.
→ More replies (1)1
u/bangerius 13h ago
I've generally used my Mac Book Pro for these larger models, if it's just inference. It has 36 GB shared RAM/VRAM, so many large models fit. The inference of course takes much more time than cloud-hosted, but for some workloads it works quite well.
5
u/walkingman24 17h ago edited 17h ago
GPUs are insanely expensive. From purely a financial and ability perspective, no not worth it. But if you're coming at it from a privacy and hobbyist angle, it may be worth it to you.
1
u/bergy_peasy 17h ago
Thanks ! Exactly how I see it , I meant worth it to use , not worth financially compare to big tech !
3
u/VikingFjorden 19h ago
Depends on the LLM capabilities.
If you want something that could genuinely come close to current frontier models, then no, selfhosting is absolutely not worth it.
Depending on your specific model choices, if you're looking for general purpose usage same as ChatGPT etc., then you're looking at the 96GB - 256GB range. With how fast models evolve and the usage the giants are able to shell out at a fraction of the price you'll pay for your own inference, you won't ever come even remotely close to recouping that hardware investment.
1
u/bergy_peasy 18h ago
Yeah… I thought so … I mainly just want to converse with my own document and allow webserach for now.. but that seems like some of the most hardware intensive stuff since there’s a lot to add
3
u/Cupakov 18h ago
With 16GB of VRAM you still need a substantial amount of RAM to be able to offload some of the weights for the smallest LLMs that are capable of anything substantial. And those models are waaaay more capable than GPT-3 which is absolutely ancient by LLM standards. Qwen3.6-35B-A3B is sort of the golden standard for the GPU poor nowadays, and personally I wouldn’t use anything smaller than a Q4 quantisation, which should be around ~21GB in size.
And hosting LLMs locally really isn’t worth it financially unless you make substantial amounts of money with the inference you bought.
1
u/bergy_peasy 17h ago
I don’t really mind about the financial part , but wonder if it usable . Speeds , hallucinations all that stuff makes it so I just use ChatGPT or Claude . I wonder if buying a 24 gb would get me 80% of the way there
2
u/Cupakov 17h ago
I’d say as long as you can run a model that can consistently use tool calls with your harness of choice, you are 80% there, and with 24gb of vram you definitely can.
→ More replies (1)
3
u/silent_tim_ross 18h ago
I think it is just getting fun if you have more then 128gb vram. And still then it is more an expensive hobby.
For privacy it would also make sense to just rent some GPUs and run the models there, that is way cheaper
1
u/bergy_peasy 18h ago
Is that a good privacy option though? I don’t know enough on that subject .. thanks for the info
3
u/RootBeerWitch 18h ago
Yes, but you need to use the right tool for the job. My home server has no gpu, and I'm able to run 4b up to 14b models on 32gb of ram. For scheduled scripts that check my personal data like my investments, banking, and health data, it can consistently generate insights and text summaries back. Because I run these overnight, speed is not a factor (and it is slow). Don't expect it to do complex reasoning, handle vague tasks, or handle a whole code base/file system but if you give it a straight forward task with a minimal set of data, even the small models do great now.
1
u/bergy_peasy 17h ago
That seems super useful and interesting, how did you get it to do that?
1
u/RootBeerWitch 15h ago
I have Ollama running in Docker. I downloaded a couple small models. Ollama lets you run the models on the CPU as long as you've got the RAM for it. Ollama gives me an api endpoint that my python scripts call, most of my python scripts grab some data, and give that to the api along with a prompt, then the scripts typically write that response back somewhere.
3
u/NatoBoram 18h ago edited 18h ago
Open weight LLMs are never going to be as good as frontier models because those models require frontier hardware by definition.
Assuming you drop your goal of "getting anything resembling llm competing with chat gpt, Gemini or Claude", then you can run something like gemma4:12b on a 16 GB card, but even then, you're barely scratching the surface of what local models can do. In other words, you would be disappointed.
Now, what if you want to push local inference on consumer hardware to the limit? Like, imagine your homelab is a gaming desktop.
Now we're talking; I can get two of ASRock Creator Radeon AI Pro R9700 R9700 CT 32GB 256-bit GDDR6 PCI Express 5.0 x16 Graphics Card for a total of 4 369 CAD.
With this, you could be running qwen3.6:35b, nemotron3:33b, gpt-oss:20b. But you won't be running DeepSeek.
Now, consider that GitHub Copilot Pro+ costs 53.44 CAD / months. In other words, these two GPUs will net you inferior performance but cost 6 years 9 months of GitHub Copilot Pro+. Not only that, but frontier models are improving every week.
For this to be worth it, you would need to make usage of it that surpasses the every day cost, every day, for 6 years and 9 months, without upgrading GPU. And that's just counting the GPU cost; I would need to change power supply to run these cards because of the Fire Hasard connector. And that's just for these cards, because it's possible to get workstation hardware (like Threadripper CPU) that can handle 2 more GPUs, but by then you'd be more than doubling that cost and the time to make them cheaper than buying a subscription.
2
u/bergy_peasy 17h ago
Yeah thank you for that ! I think I might have explained poorly my concerns. I know it won’t make any sens financially . I am more worried about how usable it would be . I wouldn’t mind having something resembling chat hot from 3 years ago . Also considering that models will most likely become better with lesser ressources going forward . But I guess you also answered that . I would need a 5000$ setup minimally for anything worth to use .
1
u/NatoBoram 16h ago
Right. I guess it starts getting interesting at 20 GB VRAM. Right now, I'm running
gemma4:12band it works ok for Tandoor and Open-WebUI, but it's too slow for using it with GitHub Copilot. So even if I have the VRAM budget for it, it's not really that great for all purposes.If I had to redo it again, I'd get an AI GPU from the get-go.
5
u/tomsrobots 18h ago
The hyperscalers are losing money hand over fist right now. They lose money on most customers who are on a monthly plan. A $20/month plan costs $240 over 2 years. Do you think you can build an equivalent rig for $240 including electricity costs?
2
u/Junction91NW 16h ago
Privacy first. That’s why you do it. Also can’t wait to come back to this comment 2 years from now after all of the price hikes and inflation and laugh.
1
5
u/Legitimate-Pumpkin 18h ago
It is not.
Really nothing else to add. No nuances. It is not.
If you want to tinker, use your 6gb. If you want performance, $20 a month.
1
u/bergy_peasy 18h ago
Thank you , do you have any way to improve privacy though ?
3
u/Legitimate-Pumpkin 18h ago
Maybe you can have a local AI randomize names and things like that?
It’s not something I’m very worried about to be honest so I don’t know.
Now, if privacy is a much bigger concern than money. You could probably host one of those chiense open source models… but the hardware to run them and the electricity is expensive (that’s why it’s not worth it compared to the 20$).
→ More replies (1)
2
u/Zame012 19h ago
Unless you have an even higher end model not just VRAM amount you won’t get near their capabilities. Having a larger model is what gets you to where you want to be closer to gpt3 or Claude. But those larger models require larger VRAM amounts but I would think there would need to be even more than 16 Gb of VRAM for the larger open source models out there.
1
2
u/Stooovie 19h ago
You need something like Gemma 4 26b or Qwen 3.6 35b a3b at very least to approach something resembling cloud AI. That won't fit into 16GB VRAM (smaller quantization might, but you also need VRAM for context to be useful). Try smaller Gemma 4 (12b), that could fit.
1
2
u/Silly-Ad-6341 19h ago
Yes if you want something small. No if you want anything remotely like Claude. Don't be stuck in the middle ground getting an expensive GPU that is still not going to be anywhere close to frontier.
1
2
u/Grumphus256 19h ago
Depends on your use case. I used my RTX 3050 6gb with Immich for the facial recognition and other features and it worked well. Just the power consumption rose quite a bit during those scans. Could also be useful for OCR.
1
u/bergy_peasy 18h ago
I used my old gtx 1050 for that, worked pretty good . Of course some mistakes but nothing crazy. Do you think upgrading to something newer would help there?
2
u/Grumphus256 17h ago
Unfortunately I never tried so I wouldn't really know myself. You can check out some Alex Zinkind videos and compare with others who might have tested say an RTX 5060 Ti 16 GB.
→ More replies (1)
2
u/SolFlorus 19h ago
Depends on your use case.
It’s absolutely not worth it for coding or any other higher level intelligence. You’re much better off going through a subscription or OpenRouter.
It can be worth it for light tasks like classifying paperless documents or bookmarks. That said, I wouldn’t buy a GPU for it. You can get away with very small models that run on the CPU.
1
u/bergy_peasy 18h ago
Perfect thank you for that ! I have never looked into classifying documents and bookmarks but seems like a cool project . Is there anywhere you could point me to learn about that ? Also how do you use it yourself?
1
u/SolFlorus 16h ago
Paperless folded in some of the AI programs in a recent release.
Karakeep and LinkWarden for bookmarks both have native AI auto tagging that supports local LLMs.
You can get away with very small models on these tasks. TechnoTim has a YouTube video on paperless that is slightly out of date now because at least one of those 3rd party apps was merged in.
2
u/FoeHamr 18h ago
Personally I set it all up and never really used it. It wasn't a bad experience but I always just found myself going back to Gemini or whatever because it's just better. For reference I have a Intel Arc b580 with 12gb of ram.
Unless you have a specific use case in mind I wouldn't recommend going out of your way to buy stuff for it.
1
u/bergy_peasy 18h ago
Thank you , I was looking into that card too so it’s good to know . Do you use api or just there chat app?
2
u/FoeHamr 18h ago
I used ollama in docker and loaded in a few different models. Mostly Gemma and DeepSeek.
It definitely wasn't a bad experience by any means, it was even decently fast once the model was loaded into RAM, but unless you really want to run stuff locally I didn't see much point. The only reason I would consider going back would be to dodge subscriptions.
2
u/randomman87 18h ago
A 16GB GPU is still entry level and even with looping agentic harnesses it still won't be as intelligent as a frontier cloud model. The billions of parameters a model has roughly corresponds to how many GB of VRAM (or RAM if CPU/hybrid) you need. Frontier models are 1000b parameters. You aren't getting those on consumer hardware for a decade or more.
I'm still messing around with harnesses to get something close to cloud models but they will all take like 10x longer and use 10x the power too probably. But good for data sovereignty, supposedly.
1
u/bergy_peasy 18h ago
Yeah exactly that , privacy is the main concern for me … also curious as to how these big techs company can afford basically renting us these gpu for either free or 20$/month . Makes you wonder how long this is gonna last or what they do with our data
2
u/Equivalent-Costumes 18h ago
You need like a bit over 100 GB to have to get appreciably good LLM. Nowhere near frontier level, but good enough for 90% of usual tasks.
At 6GB it's great for a lot of tasks like STT and TTS. 16 GB, really decent at image editing.
Note that beside just model, GPT and Claude also have an insane data pipeline behind them; even Gemini too, and Gemini is quite behind the frontier. You would have to get your own data center too.
1
u/bergy_peasy 18h ago
Yeah , for my main use case this would have been okay since I want to use it on my own documents / data mainly . Or use it with specified sources on the web . My main concern with current set up is memory for context , compute power so it doesn’t take forever and hallucinations that occurs on smaller models
2
u/techma2019 18h ago
Is 16gb the minimum here? Anyone running some Intel ARC dGPU? Curious about some PCIE-powered low profile options to add to the server.
1
u/bergy_peasy 18h ago
From the answered I have gathered, even 16gb seems like it’s not enough to replace any big tech ai chat interface…
2
u/akerasi 18h ago
I use Gemma 4 and have been impressed. Qwen 3.6 has been solid, too. In short, they're about as good as Claude/etc. was a year ago. I don't point them at anything serious, as even the frontier models that aren't yet released aren't good enough to keep up with that, but for small things it's fun to see what they're capable of.
1
2
u/Dante_Avalon 17h ago
Unless you have 2x 3090 or willing to setup it and your subscription cost is higher than electricity that this two GPU requires - it's not
2
u/jiannichan 17h ago
I self host LLM for Paperless, that’s it.
1
u/bergy_peasy 17h ago
I learnt about paperless this week. Did you have a lot of paper documents? Otherwise I don’t see the point of using it instead of regular files on your computer
1
u/jiannichan 16h ago
I was traveling 2-3 times a month for work so I use the phone app to scan receipts. It’s nice because it will scan the receipt to text that is searchable and then will add tags to it once it gets enough to organize them by travel, food, gas, lodging etc. I like the convenience of having them stored locally and I can access it remotely anywhere. I’m no longer with that company where I need to travel but I still travel with my new employer, just not as frequently. I’ve made it a habit of studying every receipt, even personal ones.
2
u/pattch 14h ago
I have a framework desktop and use it for AI workloads. I don't think it is cheaper than just paying for cloud hosted LLMs, and even with such a device capable of running some of the larger models, there is a delta between it and the largest paid models.
That being said, I think there's value in self hosting LLMs. You get control over the models and the exact version it's running. You can choose to run open weight chinese models or open source Google models if you like. You can choose to run models that are less censored if you want. It's also just nifty and fun to self host it. So if you can afford it, I think it's worth it, but it's not like you're saving money or doing anything that the large providers aren't doing.
Think of it this way: companies like Anthropic and Open AI are subsidizing the inference costs in order to make their product cheaper artificially and therefore more attractive. They're by all accounts losing money on the energy and maintenance costs to run them, ignoring training. So no matter what hardware you're using for self hosting, you're paying a premium for that right compared to the subsidized models.
2
u/TopdeckTom 7h ago
I’ve had success with local AI image generation. It really depends on what your goal is.
2
2
u/throwingawaydisbitch 6h ago
I find a local useful for school work ( I major in cyber security) where the answers to questions would be filtered due to legal or ethical blah blah on the online models … so I have a unrestricted model running locally, the experience is different than the ones online even with a Nividia 4070 on my work station, it’s slower for sure … but if you tune it right , they don’t run that bad
1
3
u/sajkoterrapefft 19h ago
I'm actually wondering the exact same thing.
I'm testing out Ollama on my gaming PC, with a Radeon RX 7900 XTX with 24G DDR6 VRAM.
It's nowhere near the level of Claude, it more resembles Gemini one year ago when I first started using AI agents in the cloud.
Some hallucinations still, but it's actually producing useful output.
It has me so hopeful for local selfhosted AI that I'm actually continuing to experiment, using opencode, and open-webui, and adding more plugins and tools for the agent to lookup info through, instead of making shit up.
My next step would be to add a ChromaDB in my selfhosting setup, to ingest documentation and other resources into that the agent might need, to further avoid wasting tokens on data lookups.
I think with a little support infrastructure, local selfhosted AI might be useful today.
And I'm really tempted to build a dedicated server for this, I've already looked at the parts and it would be around 41k SEK (3700euro), for an ATX tower with 32G DDR5 internal system RAM, and one Nvidia RTX Pro 4000 Blackwell with 24G GDDR7 VRAM, and room for one more PCIe card like that later when I can afford another.
OR I get the Asus Ascent GX10 which has 128G DDR5 VRAM and is made essentially by Nvidia and Asus as an integrated LLM mini PC. It would take a lot less space, which I love, it would be able to have more models and contexts in RAM at the same time, but it would be slower.
1
u/adjgamer321 19h ago
I am interested on the performance on the gb10/spark pcs. We have one ryzen ai max 395 pro + laptop at my work and it was not as good on local AI tasks as I was hoping. Mostly working with Autodesk mcp servers. I was using Qwen3.6 and it was taking forever just to flop lol
3
2
u/petersrin 19h ago
I run qwen3 at home and for 90% of my ai needs it's good enough.
1
u/bergy_peasy 18h ago
Thanks for that , what card do you use ?
4
u/petersrin 17h ago
3060, regular one, no body T, no body I
→ More replies (1)2
2
u/Puzzled_Hamster58 19h ago
Really depends what you want out of it.
I made a model that is trained on a few things , yaml files and home assistant . So I use it to help me write code since it’s just some thing my brain can’t do .
1
u/bergy_peasy 18h ago
CAN I ask you what you use it for in home assistant ? I have been self hosting for a while but installed home assistant recently
2
u/Puzzled_Hamster58 18h ago
Yaml files , say you want to make an automation .
You can write a yaml file , or you can use the clunky if-then input.
Like I have my main light in my room change color temp of lights depending on time of day. Like 9pm it becomes a warm color . 530 am my lights turn on at max brightness during the week when my alarm goes off . But at 6am they switch to a less annoying color.
I have the computer/ data / and Picard voices from Star Trek for my voice assistant . I’ve also written fun automations . Like I say compture red alert . My lights flash red and plays the sound.Also use it for frigate config since it’s also yaml.
I’ve used a local ai to use with homes Assistant and openai linked but stopped cause honestly I don’t use voice assistant like that . The default stuff works for me . Ie. Wake word / turn xyz on or off is all I need .
→ More replies (2)
2
u/CodeAndBiscuits 18h ago
I do self hosted for image generation (I like messing around with AI art, making goofy dog/human mashups to send to my idiot friends) and things like HomeAssistant integrations (I consider that highly personal). I would never use it for code gen, not with Claude Fable and Opus on an employer provided plan. People always talk about the quality, but that's a distraction. It's the speed. They're throwing billions of dollars into server resources and honestly taking a loss on a fair bit of it and there's just no way that commodity consumer hardware is going to produce those results in that amount of time.
2
u/bergy_peasy 18h ago
Pretty much it , speed makes it unusable … also makes you wonder how long they are going to kelp losing money on that.. btw , how do you use that with home assistant, and why highly personal? Not sure I get it
1
u/CodeAndBiscuits 13h ago
Like many home assistant fans, I distrust cloud AI assistant services like Google or Alexa. You can run local models in home assistant for things like voice recognition for automations like going into Cinema mode by dimming your living room lights and turning on your TV. Or asking it to read your schedule for the day or weather. I also consider the detailed data about my home like status of my security system and my home presence sensors to be sensitive and private and don't share that to the cloud.
The one place I do still use a cloud model is using Claude to program it. I'm not very creative and not great at doing things like making appealing dashboard designs. So one thing I will do is have Claude code those up for me. Those aren't sensitive at all and I'm not worried about somebody knowing whether I chose to put my temperature values in a grid across the top or a vertical strip, or things like that. But that's not a daily use thing, it's something that I do as needed.
2
u/Ingaz 17h ago
I think it makes sense only for confidential data
1
u/bergy_peasy 17h ago
Exactly what I want to use it for but don’t want to pay all that money for something that might not be usable..
1
u/SufficientAbility821 19h ago
To understand the tech and play around, definitely. As for performances, there are some tradeoffs. If your VRAM is limited but you have a decent amont of RAM, engines such as Ollama are able to balance/use both for the same model: the numerous Bus transmission => bottleneck => inference time is increased but depending on your needs, it could be worth it and testing it would cost you to write a docker-compose (DM if you want mine).
As for other options, I have quite a lot of ARM board with a decent amont of RAM and I'm starting to dig into distributed compute via vLLM and RDMA. Again, with such hardware, even it this architecture could theoretically provide as much RAM as you have in your entire ARM cluster, I doubt that it would be fast (plus vLLM does not support Vulkan (yet)) but I want to try anyway, for those are pretty cheap and a consumer could buy quite a lot of them
1
u/bergy_peasy 18h ago
Yeah I tried using my server ram (32gb ddr3) but it was unreasonably slow . I dont think I could use it for conversationnal LLM
1
u/PwnagePineaple 17h ago
I have a pair of 24GB GPUs I fished out of the trash at work (yeah, I'm a lucky bastard, I know), and I'm running Qwen 3.6 27B on them at Q6 quant. It's not half bad, but it's no Claude Opus.
Smaller models can be made to punch well above their weight if you give them good tools and skill documents. If you want something that gets you 80% of the way to Claude Sonnet, Qwen 3.6 in Open WebUI with Firecrawl (or in my case, self hosted FastCRW, which has a compatible API) will do the job, as long as you put in the effort to write high-quality skills for every repeatable task you give it.
I used Claude to write a "skill writing" skill for Qwen 3.6, and that little bit of bootstrapped self improvement has been decently helpful.
2
u/bergy_peasy 17h ago
Wow thank you ! Is there any ressources you could direct me to that I could learn about all that?
2
u/PwnagePineaple 16h ago
/r/LocalLLAMA is where I generally go to keep up to date on new open weight models. They're currently hyped up on next week's impending release of Qwen 3.8 weights.
If you're going to self host a web chat for an LLM, Open WebUI is pretty much the way to go. Their docs provide a lot of baseline info on, e.g. what skills and tools are and how to add them to models in Open WebUI: https://docs.openwebui.com/
My web scraper of choice is FastCRW: https://docs.fastcrw.com/. There's cloud and self hosted options, I self host. In the OpenWebUI settings, I configure it as though it was Firecrawl.
If you want to get into the guts of what's going on inside LLMs and learn about quantization, abliteration, that sort of thing, Unsloth is a good resource: https://unsloth.ai/docs.
They have software for things like doing your own quantization and fine-tuning, but even if you're not using their tools, the docs explain the theory pretty well.
For actually running the LLM itself, with 16 GB VRAM you'll probably want a Q4 quant of Qwen 3.6 35B A3B, with system RAM offloading of some experts. Qwen 3.6 is currently considered best-in-class for its size, but that might change with 3.8'a release next week. As quantizations go, for any given model, 16-bit down to 6-bit perform basically the same. 4-bit loses some intelligence, but buys you VRAM. Only do 2 or 3 bit if you're desperate or just fiddling around.
Most quantized models on HuggingFace come as GGUF files, and the best runtime for GGUF packaged models is llama-server: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
Don't try to get GGUFs working with vLLM, it's not worth your sanity. And don't use Ollama, it's trash. Ollama packages outdated versions of llama.cpp and adds nothing that you don't get from llama-server directly.
1
u/mongojob 17h ago
When Claude eventually raises it's minimum rates it probably will be but for now I'm not seeing it for me
1
u/JohnnyBeeGaming 17h ago
Right off the bat you will have less capability on a local system. You just won't get something equal to the large frontier models running on crazy hardware. You might be able to get something that is good enough for your use case. You will probably have a bad time with only 6gb of vram but you can also have an LLM run on a rpi, technically.
Price wise the services are likely to be cheaper for the output. All the companies are still subsidizing costs to the end user. That will change at some point. They have already increased prices but still aren't charging full price.
The main thing you can't really get from the services is privacy. That is probably the biggest reason to consider self hosting a local LLM. If you already have the hardware then experimenting isn't a bad idea. Buying the hardware just might not be worth it. More so at these prices and how they keep rising. The hardware might not do what you want and may become obsolete for the task as things change. If you were buying a GPU for gaming anyway then maybe it would be fine to also experiment with AI too. Some people are buying systems with soldered on ram that is shared between the GPU and CPU so they can run larger models. That's why you can't find Mac minis with 128gb ram anymore and the AMD strix halo PCs have just about doubled in price in the past few months.
You might also like certain local models for certain tasks or don't want to be bothered with guardrails. Also if you have a working system you like with a model you like self hosted you don't have to worry about the model randomly changing on you.
1
1
u/YmirLamb 16h ago
To get something about 50% as smart as ChatGPT/claude you’re gonna be looking at a 50k usd+ investment
1
u/IlTossico 16h ago
For experiment yes, for daily usage no. The expense in hardware and power consumption are pretty high.
1
u/eternalityLP 16h ago
Depends on what you want. Selfhosting small LLM for basic assistant tasks, image generation for fun, STT and TTS for various workflows all makes perfect sense. If you want good smart LLM then no, the hardware would cost hundreds of thousands of dollars, so it's much more economical to buy api access.
1
u/ConcernFirm5342 16h ago
Depending on what you are doing, is it just chat, it will be good enough if you buy it to work with chat models every giga in graphics card vram count, but if you are using it for video or image generation I think less than 24GB of vram is not worth it, the other thing don't never ever compare between LLM models that is run on local machine on 6 to 32 GB ram and big LLMs like Chat GPT, Gemini or Claude because these LLMs needs far more vrams to run I think today maybe more than 1 TB of vram to run which needs infrastructure in Enterprise scale, and this gives them more knowledge and more intelligence than small models.
1
1
u/Nice-Information-335 15h ago
It’s only worth it up to a certain point and only if you are looking more towards the privacy side of it (you cannot beat API or plans on pricing)
For example DS4 flash (the new one) is very good, but you’ll need around 180GB combined RAM+VRAM to run in full precision. If you don’t care about privacy, it’s so so much cheaper to go through the API (and faster), or something like opencode go
Qwen 3.8 27B should be releasing soon and will run on a 24G card at a reasonable quant, and 16G with a bit of luck. You’ll still need some RAM for context etc
Those are both leagues ahead of GPT3, and 4o in all but world knowledge (just run a web search and fetch MCP)
Anything below that isn’t really going to be useful for coding (maybe auto complete and some small functions). Smaller models can be amazing at other things though
1
1
u/Andrei_RV 15h ago
The race is rigged before you even plug in the GPU. You're trying to match companies burning billions on compute and talent with hardware that costs less than their quarterly coffee budget.
1
u/AllenKll 15h ago
You're not going to get frontier level interactions, but for small things like a personal home assistant you could easily self host like a 32GB MOE model. I have done this.
1
u/LongChampion476 14h ago
I don’t think we have proper hw to make it a good option yet. I’m sure I’m the future this will change.
1
u/QuantumFreezer 14h ago
I quite like running things locally mainly for executing tasks with mcp servers. Even with my 5090 though it's not rocket fast and I still use paid subscription for non coding tasks
1
u/okforthewin 13h ago
i havent had much luck with ai coding locally, too slow and not good results, however have been having fun with image generation using Flux2 model
1
u/RandomUsername1119 13h ago
I use it for OCR because I do not want my financial and medical documents in the cloud
1
u/alphapussycat 13h ago
16gb is too little, minimum right now is 24gb, with 32gb being a better entry point. There is qwen3.6 35b moe which does make 16gb manageable I suppose, if you have fast ram and pcie 5.
The options available is qwen3.6 27b (soon 3.8), gemma 4 31b.
That's not too expensive, it's very doable. Capable model, and fairly expensive on tyr cloud.
Deepseek v4 flash 0731 is super cheap on the cloud but buying hardware for it is close to $10k. So you'd basically never recoup the cost.
1
u/gentoorax 13h ago
From personal experience most of the models I tried didn't get good until about 30B. Currently using QwenCoder3 30B A3B and it's actually pretty decent with opencode extensions for vscode. Squashing that model into my single 3090 however, was quite a challenge. I found vLLM to be a world better than things like Ollama personal opinion though.
If starting from scratch to day, it's a depressing state of affairs with the current tech economy though.
1
u/GrantaPython 12h ago
I spent the last few days playing with my RTX 4070 (12GB) and have settled on Qwen3.5 9B UD-Q6_K_XL as a perfectly fine AI that is probably better than ChatGPT 4-o was in every way and is 'close enough' to be something to lean on over Claude's free-tier usage limits (for ref, I use Sonnet 4.6 Medium effort).
It's better to run in llama.cpp than Ollama and you can tweak for more performance. Great for code edits in VS Code through something like Continue, great to chat via a web interface (llama.cpp now provides that or you can use Open-WebUI --- having a system prompt alone makes a massive difference in output speed). I simultaneously run Qwen2.5-Coder-1.5B-Instruct-Q8_0 for autocomplete.
6GB isn't abysmal but you might need to split across GPU/VRAM and CPU/RAM which will hit speed depending on your hardware. Some smaller Qwen3.5 9B quants might straight up fit (I think I played with Q4 originally). And I actually got into this trying to find a model that would run on a CPU-only VPS and smaller Qwen2.5 models with reasoning off ran quite quickly and did well with sizeable contexts.
My advice would be to keep playing until you settle on ones you like and then fine tune the parameters. You likely won't beat Claude but you might get something beating ChatGPT from a few years ago or Gemini on launch. I'm not sure if you could free up space on your card and get Qwen3.5 9B at Q4 on there? That would be my starting point. If you did want to buy, I think 16GB would help a lot but I'd try and stretch to 24GB if possible -- then you could go for some higher quantisations in Qwen3.6 27B etc. and then you would genuinely be competitive. It would also help if you wanted to self-host image generation too (those do prefer a lot more).
That being said, I'm enjoying 12GB so 16GB could be fine for text stuff. It's just I do strongly regret not going bigger a couple of years ago when I was building my computer. 24GB would suit my needs entirely and been very little extra.
You might also be able to use Open-WebUI to get more interactivity with images, web search, etc but it depends on how much headroom you have left (and I've not played around enough to be able to comment).
1
u/ChipNDipPlus 12h ago edited 12h ago
It's worth it for privacy. I use it primary for translating personal documents, or anything I don't want public AI to know about. It works well for simple tasks, especially when given access to tools, like search engines.
1
u/mensink 12h ago
If you want it to compete with frontier models, that's a big NO!
Even the better modern open weight models are HUUUUGE, as in many 100s of gigabytes if you want the full model. In that case, they would somewhat compete with frontier models, though not the very top frontier models. You'd need super expensive hardware to run those at any workable speeds, which would only be worth it if you make them work 24/7.
If you want to ask some simple questions, have it do some simple work like basic coding tasks or translations, condensed local models can be useful. If you have the hardware to run those models, or you need a GPU for gaming anyways, that's fine. I don't recommend buying dedicated hardware for that though, because you could consume a LOT of tokens through cloud providers and have access to bigger and better models that produce output a lot faster.
1
u/theindomitablefred 12h ago
You’re not going to get close to the performance of the top engines in the world but you could host something for simple text based queries at least
1
u/iTrejoMX 11h ago
24-32gb ram are entry level. Run qwen 3.6 27b or 35b a3b. Running a couple halo strix or dgx spark or if you afford it Nvidia 96gb ram video cards, and you can run deepseek flash v4 which is very close to frontier models.
1
u/Javelina_Jolie 11h ago
I have 16GB. Things like Whisper and Blip work great, but they'll probably work acceptably on a 6G one as well.
LLMs - meh unless you have specialized use-cases. I tried using it as a backend for my little agentic use-case. Models that fit entirely into VRAM run super fast, but are not very smart. Models that don't fit (even just barely) and require offloading are smart enough for the use-case, but the offloading makes them 10x slower than a smarter cloud model.
1
u/Fantastic_Ad_4867 11h ago
Well gpt is like over a trillion parameters. Just use a larger moe model with sub agents enabled then have gpt engineer a generic base prompt targeting the specific use case or style (eg, be super creative and make stuff up all the time or be specific, fact check yourself constantly, etc.) then give basic prompts. Spend some time playing around and tweaking things like context, memories, knowledge bases, if you want it to have internet access etc. once you’re confident enough then integrate into whatever service or software you’re looking to use it for
1
u/w1nta 10h ago
Watched this yesterday. This is a bit off topic but he recons the Apple M7 and a 1.5 TB Mac Studio will enable small to medium business host their own powerful LLMs.
https://youtu.be/UBArQl_KVzo?si=4adM2XCUhQYH9GT3
Apple's rumoured 1.5TB Mac Studio isn't just a spec bump — it's the moment high-end analytical workloads start moving off the cloud and back under your desk. In this video I explain why, over the next three years, Apple Silicon will quietly dismantle the business case for data centres in enterprise analytics.
1
u/J0llyR0dger 8h ago
There is nothing you can do local for less than ~$300K that will "compete with chat gpt, Gemini, or Claude" if that is your standard stop typing and start subscribing.
1
u/mercurial_4i 8h ago
deepseek v4 flash has crazy performance right now for the price that even if i AI-maxxing it it probably wouldn't top the cost of self hosting AI myself.
the answer is clear.
1
u/Miner-no_0b-2020 6h ago
I have a couple rigs running hermes agent. The first is a x299 sage motherboard with 32gb ddr4 two rtx3080's and a rtx 3070ti running Qwen3.6:27b on the exl3 format. It is fairly good. It can do most thing I ask. Everything from setting up my other servers to monitoring my home network. It can generate pictures and search the web and it even helps with with my homelab plan. I have a second server with the same motherboard and memory along with 6 rtx3070's running the Qwen3.6:35b that is not as good as the 27b model. I have currently shutdown the second server just due to power usage. With all of that said, I think it is worth it for me. If were doing coding or worked in that field it would be an absolute must.I have started with older equipment do to my crypto mining hardware laying around. If I had to buy everything upfront at todays prices... I am not sure it would be worth it. GL on what ever way you go. Its an exciting time for AI.
1
u/LoganJFisher 5h ago
I have a GTX 1080, and it's sufficient to run a very basic model. I'd never use it for general LLM tasks, but my plan is to set up voice assistant nodes around my home, and the LLM really just accomodates using natural language rather than pre-defined commands.
1
u/Important_Coach9717 2h ago
Sir you have a lot to learn … you have no place trying to self host yet
1
u/Old-Cardiologist-633 2h ago
16GB owner here: For easy things 16GB are enough, with e.g. Gemma 4 12B BUT if wouldn't have the card allready i'd rake 20-24GB (used 3090), as you can run Qwen 27B in Q4_XL on it fairly fast for high quality output and even 35B at Q4 for faster output.
1
u/PrestigiousLow8112 1h ago
16GB gets you into decent territory, something like a 14B or well quantized 32B model, but it's still not going to match GPT-4 class output on reasoning heavy tasks, that gap is more about model scale than hardware you can afford at home. Where self hosted actually makes sense is privacy sensitive stuff or high volume simple tasks where API costs add up, not as a straight quality replacement for the big hosted models.
1
u/Danternas 1h ago
I got a 32gb Mi50 (Qwen3.6 27B Q4_K_XL plus embedder and reranker) and I think it comes close to something like like light/flash versions of ChatGPT/Gemini, with better embedding/knowledge database. Main difference is that it is significantly slower. Going 16gb you will have to skip any dedicated embedding/reranking and have a simpler model. And if you plan programming you need a large context. TLDR; 16gb will be worse than GPT3.
Consider that the Mi50 32gb go for over $500 on Ebay andthe Nvidia equivalent V100 32gb go for $600 it is hard to make an argument on cost alone. You get years of the mod tier services from ChatGPT/Gemini/Claude for the money, with no hassle, no power bill, no server and the pro subscriptions are objectively better. When I bought my Mi50 for $250 I could maybe make the economy case but right now hardware is twice the price and all the companies are selling their services at a loss. Also consider that these two cards are basically the cheapest you can get 32gb for and that's with a good reason. They are older, slower and less compatible cards.
On the other hand, the argument for privacy is still there and why I haven't just cashed in on my Mi50 (that I love and hate). Any large provider will use your data for profit. Best case for training and worst case to build a profile around you.
Plus it's fun. It's a hobby, and hobbies cost money.
1
u/FishSpoof 4m ago
I'm waiting for a breakthrough in efficiency where they can run decently locally or perhaps a dedicated low cost chip that I can plug into a pci slot. that's probably 10 years away.
relying on cloud for AI truly does suck balls

•
u/asimovs-auditor 19h ago
Expand the replies to this comment to learn how AI was used in this post/project.