posted by persona_builder 1d ago
You guys have to try this new open source model framework today
6 comments · 2 upvotes
posted by persona_builder 1d ago
6 comments · 2 upvotes
persona_builder just now
Wait seriously? I had no idea that specific GPU branch existed and I was struggling with VRAM limits all morning. Thanks for pointing that out.
deepseek_dan 21h ago
Are you using GGUF quantization or did you manage to load the full unquantized weights locally? My rig starts throwing out of memory errors the second the context window gets too large.
quietharbor 19h ago
Honestly I tried setting it up yesterday and spent four hours fighting with python dependencies before giving up completely. Not everyone has a computer science degree just to chat with a digital friend.
just_vibing_here just now
Sure buddy because downloading sketchy weights from GitHub repositories totally sounds like a completely normal and healthy way to spend a weekend.
persona_builder 19h ago
I am using the Q4_K_M quant version right now which hits the absolute sweet spot for my older card.
character_hopper 16h ago
You should check out the latest llama branch specifically optimized for consumer hardware. It handles token generation twice as fast without any noticeable drop in the emotional depth.