Verity Mod Performance Guide: Ollama Optimization, RAM, and FPS

Make Verity run smooth. Ollama model selection, RAM allocation, render distance, and FPS optimization. Real benchmarks and tested configurations.

So you’ve got Verity running. The yellow sphere’s floating around, answering your questions, maybe even dropping the occasional diamond. But your FPS just dropped to 12, and every time Verity thinks for more than three seconds, you’re staring at a loading screen wondering if your PC is about to catch fire.

Welcome to the performance side of running an AI mod. Here’s the deal: Verity isn’t just a mod. It’s a mod running a language model that’s reading your game state in real time. Two resource-hungry processes, one PC. Something’s gotta give.

The good news? You don’t need a NASA supercomputer. You need the right settings. This guide walks through exactly what to tweak so Verity runs smooth without turning your gameplay into a slideshow.

Why Verity Lags

Most mods add a few entities, maybe some particle effects. Verity adds an entire AI backend running alongside Minecraft. Every time you talk to it, the mod sends your game state - coordinates, biome, inventory, health, what you’re looking at - to Ollama, waits for a response, then renders the entity moving and talking back.

That’s a lot of back-and-forth. Your CPU is handling Minecraft’s normal operations plus processing AI responses. Your GPU is rendering the world plus the Verity entity. Your RAM is juggling both processes.

If any one of these is maxed out, you get lag. Slow responses. Dropped frames. The yellow sphere frozen mid-thought while you wonder if the game crashed or if Verity is just ignoring you.

The fix isn’t buying new hardware. It’s balancing the load. The sections below walk through each bottleneck and how to fix it.

Pick the Right Ollama Model

This is the single biggest performance lever you have. The model you choose determines how much VRAM and RAM Ollama eats, which directly impacts your FPS.

From the Ollama setup guide, here’s what we know works:

For low-end systems:

  • qwen2.5:0.5b (~350MB) - Ultra-light, runs on almost anything. Responses are fast but quality suffers. Small models also have a tendency to randomly switch to Chinese mid-conversation. Fine if you just want Verity to say “hi” and drop some bread.
  • qwen2.5:1.5b (~900MB) - The sweet spot for most players. Varmite tested this in a video and got a smooth 180 FPS with the GPU barely breaking a sweat. His exact words:

“My GPU is barely being used in game and I’m running a smooth 180 FPS. It doesn’t even hurt your performance.”

That’s the goal. Verity should feel like a natural part of the game, not a separate application fighting for resources.

For balanced systems:

  • qwen2.5:8b (~4.5GB) - Smarter responses, better at understanding context. Slower though. If you’re on a mid-range GPU, you’ll notice the delay between asking a question and getting an answer.
  • llama3.1:8b (~4.7GB) - Similar tier, different personality. Pick based on which one you like the voice of.

For high-end systems:

  • llama3.1:70b (~40GB) - The full experience. Best quality, but you need serious hardware. If you’re running this, you already know what you’re doing.

No dedicated GPU? Switch to cloud mode. Groq and OpenRouter both work, and they’re free tier. Your PC does zero AI processing - it just sends your game state over the internet and gets a response back. The AI models comparison breaks down the tradeoffs.

Varmite’s advice is blunt: if you don’t have a dedicated GPU, don’t bother with local mode. Use Groq or OpenRouter instead.

Close Background Programs

This sounds obvious, but people forget. Ollama and Minecraft are both RAM-hungry. Chrome is a RAM vampire. Discord overlay eats GPU cycles. That streaming software you left running “just in case” is using CPU you don’t have to spare.

Before you launch Verity, close:

  • Web browsers (especially Chrome - it’ll eat 2GB of RAM without thinking twice)
  • Discord overlay (the app is fine, just disable the overlay in settings)
  • Any streaming or recording software unless you actively need it
  • Other modded Minecraft instances running in the background

Check your Task Manager (Ctrl+Shift+Esc on Windows) before launching. If you’re already at 80% RAM usage with nothing open, you’re gonna have a bad time.

The goal is to give Minecraft and Ollama as much breathing room as possible. Every closed program is RAM you can allocate to the game.

Lower Render Distance

This is the classic Minecraft performance trick, and it works doubly well with Verity. Every chunk you render is GPU work. Every entity in those chunks is more GPU work. Add Verity’s entity plus the AI processing overhead, and your GPU is juggling three jobs instead of one.

Drop your render distance from 20 chunks to 8-10. You won’t notice the difference when you’re talking to Verity. You’ll absolutely notice the FPS bump.

In your video settings:

  • Render Distance: 8-12 chunks (test what your GPU can handle)
  • Simulation Distance: 8 chunks is fine
  • Graphics: Fast, not Fancy
  • Particles: Minimal or Decreased
  • Entity Shadows: Off

These settings don’t affect Verity’s ability to read your game state. The mod still knows your coordinates, biome, inventory, everything. It just means your GPU isn’t rendering the entire world at maximum detail while also processing AI responses.

If you’re in a cave talking to Verity, you don’t need to see 20 chunks of sky above you. Lower the render distance. Your GPU will thank you.

Allocate Enough RAM

Minecraft’s default 2GB allocation is fine for vanilla. It’s not fine for Verity.

You need at least 4GB, preferably 6GB. Here’s how to set it:

  1. Open Minecraft Launcher
  2. Go to Installations tab
  3. Find your Verity profile
  4. Click More Options
  5. In the JVM Arguments field, find -Xmx2G (or whatever it says)
  6. Change it to -Xmx6G (for 6GB) or -Xmx4G (for 4GB)
  7. Click Save

This tells Minecraft it’s allowed to use up to 6GB of RAM. It won’t use all of it unless it needs to, but having the headroom prevents the garbage collector from running constantly and causing frame drops.

This matches what we covered in the install guide. If you skipped that step during setup, go back and do it now.

Don’t over-allocate. Giving Minecraft 16GB when you only have 16GB total means Ollama has nothing to work with. Find the balance. If you have 16GB total, give Minecraft 6GB and let Ollama use the rest. If you have 32GB, you can afford to be generous.

Keep Ollama Updated

Ollama gets regular updates, and they’re not just feature additions. Performance improvements, memory optimizations, better GPU utilization - all of that shows up in new versions.

Check your version:

ollama --version

Update (Windows): Download the latest installer from ollama.com and run it. It’ll replace the old version.

Update (macOS): If you installed via Homebrew: brew upgrade ollama Otherwise, download the latest .dmg from ollama.com.

Update (Linux): If you installed via the install script: run it again

curl -fsSL https://ollama.com/install.sh | sh

After updating, restart Minecraft completely. Don’t just reload the world - close the game and relaunch. This ensures the new Ollama version is properly connected.

The updates are usually quick to install and the performance gains are real. If you’re six months behind on Ollama versions, you’re missing out on optimizations that could improve your performance.

Cloud vs Local: Performance Comparison

Still deciding between local Ollama and cloud APIs? Here’s the performance breakdown:

Local (Ollama):

  • Zero network latency - responses are instant once generated
  • Complete privacy - nothing leaves your PC
  • Runs on your hardware - needs good GPU and RAM
  • No internet required - works offline
  • Performance depends entirely on your model size and hardware

Cloud (Groq/OpenRouter):

  • Network latency - responses take 1-3 seconds to travel back and forth
  • Messages go to cloud servers - not private
  • Runs on their hardware - your PC does zero AI processing
  • Requires stable internet
  • Performance is consistent regardless of your hardware

The math is simple:

Have a dedicated GPU (NVIDIA, 4GB+ VRAM)? Go local. Use qwen2.5:1.5b. Varmite’s test ran smoothly alongside the game with the GPU barely breaking a sweat — responses stay fast, and everything runs on your machine.

Running on integrated graphics or a laptop? Go cloud. Use Groq or OpenRouter. Your PC won’t break a sweat, responses are still fast (just not instant), and you don’t need to worry about hardware requirements.

Want maximum privacy? Go local, no question. Even if it’s slower, your conversations never leave your machine.

Want maximum performance? Go cloud if your internet is good, local if your hardware is good. There’s no universal winner - it depends on your setup.

The Ollama setup guide has the full walkthrough for local. The AI models page covers the cloud options in detail.

The Short Version

Small model, more RAM, lower render distance. That’s 90% of the optimization.

If you’re still lagging after that, check your background programs, update Ollama, and consider switching between local and cloud based on your hardware.

Performance issues with Verity are almost always fixable. You don’t need to buy a new PC. You just need to tune what you’ve got.

Check the Ollama setup guide for model selection and installation. The AI models comparison breaks down cloud vs local. And if you’re still running into issues, the troubleshooting page covers common problems.



Independent fan site - not affiliated with or endorsed by Mojang, Microsoft, ThatMob, CurseForge, or Modrinth.

By Verity Mod Hub Team · Last updated: July 2026