I Pushed My Local AI’s Memory to the Limit
Testing context window limits on local AI — what worked, what slowed down, and what crashed.
Educational AI demos, tools, and experimental builds by NosisTech. Published for learning and documentation only, not as production systems, implementation guidance, or professional advice.
Testing context window limits on local AI — what worked, what slowed down, and what crashed.
I set all of this up and do not fully understand everything I did. But it works, it is fast enough to be useful. That is the point of documenting it. The Hardware My machine is an Acer Predator PHN16S-71. CPU: Intel Core Ultra 9 275HX. GPU: NVIDIA RTX 5070 Ti with 12GB VRAM. RAM: … Read more
I set up my own AI server on a Hostinger VPS running Ollama, proxying Claude through LiteLLM, and eventually adding Gemini. Here’s what that actually looked like.
I didn’t want to pay for a VPS just to access my local AI setup remotely. Cloudflare Tunnel solved it for free — and it took about ten minutes to set up.
I installed Ollama and two AI models on my Acer Predator laptop. It wasn’t smooth — but it worked. Here’s an honest look at what happened and what I learned along the way.