Skip to content
View forcepusher's full-sized avatar

Block or report forcepusher

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
forcepusher/README.md

Projects done with substantial help of AI tools have "Slop" in their name. Other projects untouched by AI.

Of course "AI" is a great to get stuff done fast, but it's quite dumb and have to be very carefully guided.
Some people compare it to a parrot facerolling on the keyboard.

LLM is not even a neural network, it's an autocomplete dictionary for T9 text predictions just like in old phones.
Repeatedly tap on your phone's text predictions - this is the current state of "AI".
Now with proper expectations you're ready to start building.

BTW, if you don't want to feed money to cloud services - start with your own local LMStudio/ComfyUI machine.
All you need is 16GB GPU and 32GB RAM to start. CPU doesn't matter, it's really that cheap.
Setup takes 3 weeks of pure suffering and you're ready for a true AI future, it'll pay off in less than a year.
Our videocards now can not only run games, but write somewhat useful code. That's pretty cool right?

And if part of your job or pipeline can actually be replaced by a parrot, maybe it should be replaced.
Think of writing and updating tests. If you're blank-staring at the wall right now, you get it.
But don't let LLMs think for you or build an architecture - it's all pointless randomized garbage.


Cookbook (reliable agentic models for programming, most useful first):

Sweet spot: 32-40GB GPU VRAM + 64GB RAM
Qwen 3.8 27B ~250k - unsloth/qwen3.8-27b@q4_k_s (temp 0.1, top k 40, no rep penalty)
Gemma 4 31B ~100k - unsloth/gemma-4-31b-it@q5_k_xl (temp 0.3, top k 40, rep penalty 1.1)
Xortron ~200k - xortron.criminalcomputing.2026.27b.next@q6_k (temp 0.3, top k 40, rep penalty 1.1)

Mostly usable: 24GB GPU VRAM + 64GB RAM
Qwen 3.8 27B 100k - unsloth/qwen3.8-27b@iq3_s (temp 0.1, top k 40, no rep penalty)
Gemma 4 31B QAT 32k - unsloth/gemma-4-31b-it-qat@q4_k_xl (temp 0.3, top k 40, rep penalty 1.1)
Xortron 40k - xortron.criminalcomputing.2026.27b.next@q5_k_m (temp 0.3, top k 40, rep penalty 1.1)

Barely usable: 16GB GPU VRAM + 32GB RAM
Qwen 3.8 27B 32k - unsloth/qwen3.8-27b@iq3_xxs (temp 0.1, top k 40, no rep penalty)
Xortron 32k - xortron.criminalcomputing.2026.27b.next@iq3_xs (2 layers on CPU, Q8 KVCache, temp 0.3, top k 40)

Old videocard or laptop option, pure suffering: 8-12GB GPU VRAM + 32GB RAM
Gemma 4 12B QAT ~80k - unsloth/gemma-4-12b-it-qat@q4_k_xl (temp 0.1, top k 40, rep penalty 1.1)

--

Global settings: Min P Sampling 0.05, Top P Sampling 0.95.
This is how these settings work (yeah I know, pretty much every IT video).

If you can get anything done on a small model, get a dual-16GB-GPU setup. I use dual RTX 4000 Ada for 40GB VRAM.
Every 8GB extra VRAM is an astronomic leap in quality. 12GB model is not even close to other models.

Use OpenAI-compatible API to connect to LM Studio. The https://zed.dev/ seems to be best open-source agentic IDE.
Here are jinja templates for LM Studio and Zed. Very tedious to get right.
Put Responses MUST be terse and short. in a rule or system prompt, or use my PortableAgent ruleset.
Vision consumes a lot. Use Q8_0 or BF16 .mmproj files so you don't have to blind the model completely.

Model developers put default settings tailored for high scores in benchmarks. Those are really bad for actual work.
I use very low temperatures to avoid tool use typos/screwups, since I use LLMs mostly for routine like refactoring.
To avoid Gemma 4 thinking bugs, use "<|channel>" as your reasoning start string, not "<|channel>thought".
Disable Unified KV Cache and set Max Concurrent Prediction to 1 to save memory. Set max output tokens to 8000.


More Unity packages:

ComfyUI nodes:

Other instruments:

  • PortableAgent - Rule prompt for local LLMs like Gemma and Qwen. Read less slop and get much better results.
  • ComfyUI-SloppyInstall.bat - Simplified pip install -r "requirements.txt" for custom nodes in portable ComfyUI.
  • SloppyServer.bat - Single file local/Wi-Fi server for debugging multithreaded mobile Unity WebGL builds.

Technical articles (No AI tool ever touched this holy grail):

Pinned Loading

  1. com.bananaparty.yandexgames com.bananaparty.yandexgames Public

    Unity package. Yandex Games SDK for the WebGL platform.

    C# 71 24

  2. com.bananaparty.webutility com.bananaparty.webutility Public

    Unity package. Tools for fixing issues in the WebGL platform.

    C# 24 3

  3. FullscreenWindowTemplate FullscreenWindowTemplate Public

    Unity WebGL template that scales to the entire browser window.

    HTML 20 5

  4. com.bananaparty.yandexmetrica com.bananaparty.yandexmetrica Public

    Unity package. Yandex Metrica SDK for the WebGL platform.

    HTML 13 2

  5. com.bananaparty.behaviortree com.bananaparty.behaviortree Public

    Unity package. Fully cross-platform Behavior Tree.

    C# 52 6

  6. com.bananaparty.websocketrelay com.bananaparty.websocketrelay Public

    Unity package. Fully cross-platform WebSocket networking library.

    C# 16 1