What Am I Looking At?
A local-first visual assistant for macOS — it sees the scene, names the objects, and answers your questions. On your Mac. Nothing uploaded.
MIT License Python 3.11+ macOS Apple Silicon Local-first privacy pytest
Quick start · How it works · Stack · Docs · License
Senti — local-first visual assistant on macOS
Most visual assistants ship camera frames to a cloud API. Senti does not.
It is built for Apple Silicon so the fast path — object detection and tracking — stays at interactive rates, while the slow path — a local vision-language model via Ollama — only runs when the scene actually changes or you ask a question.
| 🎯 Live boxes & track IDs | 💬 Follow-up questions | 🔍 Object focus |
|---|---|---|
YOLO26 labels, confidence, stable #1 #2 IDs |
"What’s that connector?" "What objects do you see?" | Crop #2 phone before the VLM sees it |
| 📝 On-device OCR | 🔊 Spoken answers | 🎙️ Push-to-talk |
Ask read this or auto-read on READY |
macOS TTS (Qt or say) |
Local Whisper, Esc to cancel |
| Capability | Detail | |
|---|---|---|
| 🎥 | Live perception | YOLO26 on Apple Silicon — PyTorch mps or yolo-mlx Metal — plus ByteTrack / BoT-SORT IDs |
| 🧭 | Scene intelligence | Visual + object + spatial change detection, stability gating, best-frame selection |
| 🧠 | Local understanding | Ollama VLM, auto-analysis when the scene settles, conversational memory |
| ✂️ | Object focus | Padded crop of the most likely target before a focused question |
| 🔤 | On-device OCR | EasyOCR on demand (read this) or automatically when the scene is ready |
| 🗣️ | Speech I/O | macOS TTS and local Whisper push-to-talk |
| 🔐 | Privacy by design | No cloud uploads, no disk recordings, bounded in-memory frame buffer |
Two loops keep the UI live. The camera never waits on the language model.
Camera → detect → understand → speak
flowchart LR
Cam["📷 Camera"] --> Fast
subgraph Fast["⚡ Fast loop — 15–30 FPS"]
YOLO["🎯 YOLO26 + tracking"]
Scene["🧭 Scene change + stability"]
YOLO --> Scene
end
Fast --> Slow
subgraph Slow["🌙 Slow loop — on change or question"]
Frame["🖼️ Best-frame selection"]
OCR["🔤 Optional OCR"]
VLM["🧠 Local VLM"]
Frame --> OCR --> VLM
end
Slow --> UI["🖥️ Desktop UI + TTS"]
Mic["🎙️ Push-to-talk"] --> UI
UI --> Ask["💬 Ask / Focus / Analyze"]
Ask --> Slow
- ⚡ Fast loop — frames go to YOLO26, then tracking and scene-change detection. Target: 15–30 FPS.
- 🌙 Slow loop — when the scene reaches
READY, Senti picks the sharpest, most stable frame from a rolling buffer and sends it to the local VLM (and OCR, if enabled). Follow-ups reuse scene memory when they can.
State machine: WATCHING → SCENE_CHANGED → WAITING_FOR_STABILITY → READY
Full package map and thread model: Architecture .
Every runtime dependency, with a badge that opens its official site. Click through — these are the projects Senti stands on.
Python macOS PySide6 Qt NumPy OpenCV Ultralytics YOLO26 PyTorch MLX yolo-mlx Ollama EasyOCR faster-whisper sounddevice python-dotenv pytest
| Project | Official site | Role in Senti | |
|---|---|---|---|
| Python | python.org | Runtime (3.11+) | |
| macOS | apple.com/macos | Camera, TTS say, permissions |
|
| PySide6 / Qt | doc.qt.io/qtforpython-6 · qt.io | Native window, AVFoundation capture, TTS | |
| NumPy | numpy.org | Frame arrays | |
| OpenCV | opencv.org | Overlays, sharpness, scene diff | |
| Ultralytics YOLO26 | ultralytics.com · YOLO26 docs | Detection + tracking | |
| PyTorch | pytorch.org | Default YOLO path: MPS on Apple Silicon | |
| 🍎 | MLX / yolo-mlx | MLX · yolo-mlx | Optional native Metal detector (YOLO_RUNTIME=mlx; pip install "yolo-mlx[tracking,convert]") |
| Ollama | ollama.com | Local vision-language model | |
| 🔤 | EasyOCR | jaided.ai/easyocr · GitHub | On-device text recognition |
| 🎙️ | faster-whisper | SYSTRAN/faster-whisper | Push-to-talk transcription |
| 🎚️ | sounddevice | python-sounddevice.readthedocs.io | Microphone capture |
| ⚙️ | python-dotenv | GitHub | .env configuration |
| pytest | docs.pytest.org | Unit tests |
Pinned versions live in requirements.txt.
- 🍎 macOS on Apple Silicon (developed on M2 Pro)
- 🐍 Python 3.11 or newer
- 📷 Built-in or Continuity Camera
- 🔐 Camera permission (and microphone if voice input is on)
- 🦙 Ollama with a vision model for scene descriptions
git clone https://github.com/TeslaNeuro/Senti.git cd Senti python3 -m venv .venv source .venv/bin/activate pip install -r requirements.txt cp .env.example .env
Pull a local vision model, then launch:
ollama pull gemma4 ./scripts/run.sh
Or:
source .venv/bin/activate
python -m appOn first launch, macOS will ask for camera access. Grant it to Terminal (or your IDE) if you start Senti from the command line. Ultralytics downloads yolo26n.pt into models/ automatically (~6 MB) unless that file is already there.
Optional native Metal path: install yolo-mlx with pip install "yolo-mlx[tracking,convert]", then set YOLO_RUNTIME=mlx.
💡 If the preview is black, another app (FaceTime, Zoom, Chrome, ...) likely holds the camera. Quit it and relaunch.
| Action | What happens |
|---|---|
| 👀 Watch the preview | Live boxes, labels, confidence, track IDs (#1, #2, ...) |
⏳ Wait for READY |
Automatic VLM description of the best buffered frame |
| 💬 Type in Ask | Follow-up against scene memory, or a new VLM call when needed |
| 🎯 Focus dropdown | Analyze a specific tracked object (cropped when enabled) |
| 🔁 Analyze | Force a fresh VLM pass (bypasses cooldown) |
| 🧹 Clear | Reset scene memory and the response panel |
| 🔊 Speak / say that | Replay the current answer (when TTS is on) |
| 🎙️ Mic / Stop | Push-to-talk; Esc cancels an in-progress recording |
| 📖 read this | Run OCR on the current scene |
Status bar shows camera, YOLO device, VLM activity, FPS, and inference latency. Hover a pill for the full detail.
Step-by-step walkthrough: Usage .
Copy .env.example to .env. Important defaults:
| Variable | Default | Purpose |
|---|---|---|
CAMERA_WIDTH / CAMERA_HEIGHT |
1280 / 720 |
Capture size |
YOLO_MODEL |
yolo26n.pt |
Weights filename; stored in models/ (yolo26s.pt is more accurate) |
YOLO_RUNTIME |
auto |
auto (Ultralytics unless YOLO_DEVICE=mlx), ultralytics (PyTorch MPS), or mlx (yolo-mlx Metal) |
YOLO_DEVICE |
auto |
auto → PyTorch mps on Apple Silicon; mlx for yolo-mlx |
VLM_MODEL |
gemma4 |
Ollama vision model |
VLM_BASE_URL |
http://localhost:11434 |
Local Ollama API |
OCR_ENABLED |
false |
On-device text recognition |
TTS_ENABLED |
false |
Speak answers aloud |
VOICE_ENABLED |
false |
Whisper push-to-talk |
Optional features (OCR, TTS, voice) are off until you turn them on. Full reference: Configuration .
🎛 Optional extras
OCR_ENABLED=true TTS_ENABLED=true TTS_RUNTIME=auto VOICE_ENABLED=true VOICE_MODEL=base
TTS_RUNTIME=auto uses Qt when it exposes your voice, otherwise the macOS say command (so names like Tessa still work). First OCR or Whisper use downloads models in the background.
source .venv/bin/activate
pytest tests/ -q| Guide | Contents | |
|---|---|---|
| 🏗️ | Architecture | Layers, threads, scene states, package map |
| 🖱️ | Usage | Window tour, questions, focus, speech, voice |
| ⚙️ | Configuration | Every .env setting and sensible ranges |
| 🩺 | Troubleshooting | Camera, Qt, Ollama, OCR, TTS, Whisper |
| 🛡️ | Security | Privacy model and how to report issues |
| 📜 | License | MIT License |
Senti is local-first:
- 🚫 Camera frames are never written to disk
- 🧠 The frame buffer is bounded and in-memory only
- 💻 YOLO, OCR, Whisper, and the VLM run on-device (Ollama on localhost)
- 📡 No telemetry, no cloud uploads, no account
See SECURITY.md for the threat model and reporting.
Senti/
├── app/ Application package
│ ├── camera/ Qt / AVFoundation capture + in-memory buffer
│ ├── detection/ YOLO26 worker
│ ├── tracking/ ByteTrack / BoT-SORT monitor
│ ├── perception/ Scene change, stability, best-frame selection
│ ├── vision/ Local VLM, scheduler, object crops
│ ├── scene/ Conversational scene memory
│ ├── ocr/ EasyOCR worker
│ ├── speech/ Qt TTS + macOS say
│ ├── voice/ Push-to-talk + faster-whisper
│ └── ui/ Native desktop window
├── docs/ Architecture, usage, configuration
│ └── assets/ README graphics (logo, banner, pipeline)
├── models/ YOLO26 weights (gitignored *.pt / *.npz)
├── scripts/run.sh Create venv, install, launch
├── tests/ Unit tests
└── resources/Info.plist Camera and microphone usage strings
Senti is released under the MIT License.
Copyright (c) 2026 Arshia Keshvari.
Senti
Built on-device. Nothing leaves your Mac.