Run lightweight AI models locally on your Android phone using Termux and Ollama. One command to install, no root needed, runs completely offline.
Nothing leaves your phone unless you ask it to. The model server listens on the phone itself; momos serve --lan is the one command that opens it to your Wi-Fi.
- Checks your phone's architecture, RAM and storage
- Installs Ollama as a native Termux package — no container, no root
- Downloads a model sized for your device
- Adds a simple
momoscommand to chat, serve the web UI and manage models
Open Termux and run:
bash -c "$(curl -fsSL https://raw.githubusercontent.com/Sidharth-e/MOMOS/main/scripts/momos.sh)"The installer checks your hardware, recommends a model, installs Ollama,
downloads the model, and installs the momos command. It prints how to start
when it is done.
Larger models (14B, 32B) require too much memory and quickly crash or throttle on phones. MOMOS focuses on models under 7B parameters that fit within typical smartphone RAM limits (2GB to 8GB+):
| Model | Tag | Download | RAM Needed | Best For |
|---|---|---|---|---|
| Llama 3.2 1B | llama3.2:1b |
~1.3 GB | 2 GB+ | Fast responses, low battery use, low-end phones |
| DeepSeek R1 1.5B | deepseek-r1:1.5b |
~1.1 GB | 2 GB+ | Step-by-step reasoning and logic |
| Llama 3.2 3B | llama3.2:3b |
~2.0 GB | 4 GB+ | Everyday chat, writing, and summarization |
| Qwen 2.5 3B | qwen2.5:3b |
~1.9 GB | 4 GB+ | Math, coding, and multilingual queries |
| DeepSeek R1 7B | deepseek-r1:7b |
~4.7 GB | 6–8 GB+ | Advanced reasoning on higher-RAM devices |
| Qwen 2.5 7B | qwen2.5:7b |
~4.7 GB | 6–8 GB+ | Detailed coding and technical answers |
You can also use any model tag from the Ollama library (e.g. phi4-mini, smollm2:1.7b).
momosShows a numbered menu: chat, open the web UI, list, pull and delete models, view logs, update, or uninstall.
momos chat # Chat with your last used model
momos chat llama3.2:1b # Chat with Llama 3.2 1B
momos chat deepseek-r1:1.5b # Chat with DeepSeek R1 1.5B
momos chat qwen2.5:3b # Chat with Qwen 2.5 3Bmomos models list # Show downloaded models
momos models pull qwen2.5:3b # Download a model
momos models delete llama3.2:3b # Remove a modelmomos logs # Follow the server and web UI output (Ctrl+C to exit)Note
The Ollama server starts automatically in the background when running momos chat or managing models, and it starts on the phone only. You only need momos logs if you want to inspect server output directly.
momos serve # Start the server in the background, on the phone only
momos serve --lan # Start it exposed to your Wi-Fi
momos stop # Stop the background serverBoth momos serve and momos ui hand the terminal straight back and keep
running. Their output goes to ~/.momos/server.log and ~/.momos/ui.log, which
momos logs follows together — tail labels each line with the file it came
from, so the two are never mixed up.
Tip
A backgrounded server stops when Android puts Termux to sleep. Run
termux-wake-lock first if you are leaving the phone alone — see
Keeping Termux alive in background.
momos chat and momos models start the server for you on 127.0.0.1 and leave it running. That is deliberately private: nothing on your network can reach it. momos serve --lan is the only way to change that.
momos serve --lanThis rebinds Ollama to all interfaces and tells it which browser origins to accept, then prints the addresses:
Ollama: http://192.168.1.42:11434
OpenAI-compatible: http://192.168.1.42:11434/v1
(any non-empty key; the server ignores it)
Chat page: run 'momos ui' in this session, then open
the URL it prints (port 8080 by default)
It returns to the prompt once the server answers, so momos ui can be started
from the same session afterwards, and the phone and anything else on your Wi-Fi
can both chat with the same model.
The command only reports success after it has confirmed the server is really
reachable on the phone's Wi-Fi address. A 127.0.0.1 server looks identical to
an exposed one from inside the phone, so a bind that quietly failed would
otherwise be announced as exposure.
OpenAI-compatible API
http://<phone-ip>:11434/v1 is a drop-in OpenAI base URL, so existing SDKs and tools work unchanged:
curl http://192.168.1.42:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"llama3.2:3b","messages":[{"role":"user","content":"hi"}]}'Some clients insist on an API key. Ollama requires a non-empty one and then ignores it, so any string will do — ollama is the conventional choice.
Warning
There is no authentication. Anyone on your Wi-Fi who has the address gets the whole API, not just chat — they can list, pull and delete your models. The origin check stops other websites from reaching your phone; it does nothing about curl. Use --lan only on a network you trust, and run momos stop when you are done.
Note
Exposure is not remembered. A server started later by momos chat or momos models comes back private, and so does one after a reboot — re-run momos serve --lan. If you run it while a private server is already up, it says so rather than appearing to succeed: the bind only changes on a restart, so it tells you to momos stop first.
Tip
If you only want to chat from a browser, you do not need --lan on the phone that is serving the page. See below.
momos ui # Serve on port 8080
momos ui 9000 # Serve on a different port
momos ui stop # Stop the background web serverThe UI is served to your Wi-Fi, so you can open it from a laptop at the network address it prints:
MOMOS UI — running in the background
Phone: http://localhost:8080
Network: http://192.168.1.42:8080 <- open this on your laptop
The page is a chat interface with a model picker and a list of saved chats. Replies are rendered as Markdown — headings, lists, tables, and fenced code with the language named — by a small parser written into the page, since fetching one over the network would be the one thing on it that fails when the phone is offline. It covers what a small model actually writes and shows anything else as plain text rather than guessing.
Send starts a reply and Stop cancels one mid-generation, which matters when
a small model gets stuck in a loop. While a reply arrives the composer shows how
fast it is going, and each finished reply is stamped with the model that wrote
it. The ↻ beside the model picker re-reads the installed models. Run on the
phone, momos ui also opens the page in the phone's browser for you.
The chat list collapses to a strip of initials, or slides in over the conversation on a phone; above it is a box to filter by title.
The picker offers everything momos models would list, and each chat keeps the
model it was using — so a DeepSeek R1 reasoning thread and a small quick model
can sit side by side. A new chat starts from whichever model you used last, or
from one you pick if the page opens with none chosen. Picking a different model
mid-conversation changes who answers from the next message on, and the replies
already on screen stay as they were.
Chats are saved in the browser, not on the phone, and reopening the page brings back the one you left off in. Two consequences worth knowing:
- The list is per browser. The laptop's chats and the phone's chats are separate, and so are two browsers on the laptop. A chat begun on one does not appear on the other.
- Clearing site data clears them. A private window will not save them at all, and says so when it opens.
Keeping them on the phone instead would mean a second server process holding state, which is memory the model wants more than your chat titles do.
Reasoning is not saved with a chat — the fold is working-out you have already read, and keeping it would multiply the storage for no return.
Opened on the phone, it reaches Ollama over loopback and needs nothing else. Opened from a laptop, it needs momos serve --lan running too.
Running momos ui again while it is already serving reports the running one
rather than starting a second. momos stop does not stop it — the web server
has its own lifecycle, and momos ui stop is what ends it.
Warning
The port has no authentication — anyone on your Wi-Fi who has the address can
open the page. Stop it with momos ui stop when you are done.
Important
Both the page and the --lan flag ship with the installer, so an existing
install picks them up from momos update. Running momos ui alone on an
install that predates them will serve the old placeholder page, and
momos serve --lan will ignore the flag. Update first.
momos ui uses darkhttpd (about 1MB),
installing it on first use. If darkhttpd is unavailable it falls back to
python3 -m http.server instead.
momos update # Update MOMOS scripts and Ollama
momos uninstall # Remove Ollama, MOMOS and (if you confirm) models
momos help # Show all commandsmomos uninstall stops both background servers, removes Ollama, the launcher
and ~/.momos, and asks before deleting downloaded models.
The installers and momos update follow the ref they were installed from. Set
MOMOS_BRANCH to install or update from a branch — the value is checked before
it reaches a URL, and it is recorded so later updates keep following it:
MOMOS_BRANCH=my-branch bash -c "$(curl -fsSL https://raw.githubusercontent.com/Sidharth-e/MOMOS/my-branch/scripts/momos.sh)"- Android 7.0+
- 64-bit device (
arm64orx86_64) — Ollama's Termux package is built for 64-bit only. On a 32-bit device, see 32-bit devices - Termux (install from F-Droid or GitHub, avoid Play Store builds)
- RAM: 2GB minimum (4GB+ recommended for 3B models, 6–8GB for 7B models)
- Free Storage: 2GB minimum; ~2.5GB for a 3B model, ~6GB for a 7B model
- No root required
A 32-bit phone — or a 64-bit phone with a 32-bit Termux, which reports
armv8l — cannot install the native Ollama package. Both installers notice
this and point you at the legacy script, which sets up a Debian container via
PRoot and installs Ollama inside it instead:
bash -c "$(curl -fsSL https://raw.githubusercontent.com/Sidharth-e/MOMOS/main/scripts/legacy/proot/momos.sh)"It needs roughly 1–2GB more storage and runs more slowly, but the momos
command it installs is the same. Running scripts/setup.sh picks the right one
for the device automatically.
If you just installed Termux, run this setup script first. It updates packages, grants storage permissions, and then offers to install MOMOS — picking native or legacy for your device:
bash -c "$(curl -fsSL https://raw.githubusercontent.com/Sidharth-e/MOMOS/main/scripts/setup.sh)"Manual setup commands
pkg update && pkg upgrade -y
pkg install curl -y
termux-setup-storageThen run the Quick Install command above.
Termux (Android)
├── Ollama server (native, background)
│ └─ local model (<7B)
├── darkhttpd ──> the web UI page
└── 'momos' launcher command
32-bit devices: the same CLI, with Debian
under PRoot providing Ollama instead.
- Ollama: installed as a native Termux package and run in the background —
no container, no root. It listens on the phone only unless you pass
--lan. - Web UI: a self-contained static page, served to the phone or your Wi-Fi by darkhttpd.
- MOMOS launcher: a bash script in
$PREFIX/bin/momosthat handles starting the server, serving the page, attaching to chats, and managing models with simple arguments.
Check the log:
cat ~/.momos/install.logRun:
termux-setup-storageThen run the installer again.
The model is using more RAM than your device has available. Switch to a smaller model:
momos chat llama3.2:1b
# or
momos chat deepseek-r1:1.5bRun termux-wake-lock to keep Android from sleeping Termux during inference.
That build is 64-bit only. The installer prints the legacy command to use
instead — see 32-bit devices. scripts/setup.sh makes this
choice for you.
Two causes look identical from the browser, so check both.
The server is private. momos chat and momos models start Ollama on the phone only. Run:
momos serve --lanIf it reports that the server is already running but only on this phone, that is the answer — run momos stop and then momos serve --lan again. The bind only changes on a restart.
The browser's origin is no longer allowed. --lan pins the origin to the phone's address at the moment it starts. If the phone got a new address since (a fresh DHCP lease, or a different network), the page still loads but its requests are refused. Restart momos serve --lan to re-pin it.
The Ollama server and the web server have separate lifecycles. momos stop
ends the model server; momos ui stop ends the page.
It reads the last model you used. Pick one in Termux first:
momos chat llama3.2:1bMIT · View License



