Eight kinds of AI to put in an app that are not chatbots
August 25, 2026
* Most of these run on the device. No server, no connection.
Eight kinds

| Kind | What it does |
|---|---|
| Body, hand, and face tracking | Reads joint positions from a camera |
| Object detection | Finds what is in a photo and how many |
| Text recognition (OCR) | Pulls written text out of a photo |
| Sound classification | Tells apart sounds that aren't speech |
| Speech to text and back | Turns speech into text and text into speech |
| Semantic search (embeddings) | Finds by meaning, not matching words |
| Recommendation | Picks what this person will like |
| Anomaly detection | Flags signals that break the usual pattern |
Body, hand, and face tracking
Joint coordinates come out of a camera feed. Shoulders, elbows, finger joints, facial landmarks. MediaPipe is the common choice here.
- Good for: posture correction, counting reps, sign language, controlling things with expressions
- Needed when an augmented reality app has to anchor to a real place like "just above the shoulder"
- Runs on the device
Object detection
Draws boxes around what's in a photo and says what each one is. YOLO is the usual family.
- Good for: counting stock, reading what's in a fridge, telling pills apart
- Blocker: off-the-shelf models only know common objects. Recognizing one specific product means retraining on photos of it
Text recognition (OCR)
Pulls text out of an image.
- Good for: receipts, business cards, reading documents aloud
- Blocker: handwriting is far less accurate than print, and non-Latin scripts vary widely by model
Sound classification
Tells apart sounds that aren't speech. A different job from speech recognition.
- Good for: a baby crying, snoring, noise logs, birdsong
- Blocker: you decide what to distinguish first, then collect those sounds
Speech to text and back
Converts speech into text and reads text aloud.
- Good for: read-aloud apps, dictated journals, video captions
- Verified: a three-second Korean clip transcribed in 2.1 seconds with no GPU, no errors
- Verified: Windows ships a system voice that turns text into an audio file
Semantic search (embeddings)
Sentences become lists of numbers, so a query finds text that means the same thing even when no words match. Searching "dog food" turns up "pet nutrition."
- Good for: document search, finding similar entries, catching duplicates
- Cost: close to free
Recommendation
Looks at what someone picked before and suggests what's next.
- Blocker: it needs history. There is nothing to offer a first-time user
Anomaly detection
Learns the normal pattern and flags values that fall outside it.
- Good for: numbers that spike, stock discrepancies, early signs of equipment failure
- Blocker: enough of "normal" has to pile up before it can start
What's been verified
| Actually run | Speech recognition and speech synthesis |
| Capability checked only | The other seven |