Two years ago every app sprouted a sparkle icon and a chat window, and most of those windows are now ignored. Users have not stopped wanting AI; they have stopped wanting to talk to it. They now expect AI features in mobile apps to sit inside the screens they already use, and they notice when they do not.
What changed
Three things happened in the last eighteen months. Model costs fell far enough that running inference on every search query or every uploaded document is cheaper than the support ticket it prevents. The big consumer apps, from email to maps to banking, quietly added AI to existing features rather than adding chat, which reset what “normal” looks like. And on-device models reached the point where a phone can summarize a document without sending it anywhere, which removed the privacy objection for a whole class of features.
The cost point is worth making concrete. Embedding a search query typically costs a small fraction of a cent. A support ticket typically costs several dollars in staff time. If better search prevents even a handful of "I can't find my invoice" tickets a month, the inference bill is noise.
The result is that the bar moved. An app whose search needs the exact keyword, whose forms need to be typed, and whose long lists need to be read now feels old.
The three AI features in mobile apps users now expect
Search that understands a sentence. “Invoices from the plumber last spring” should work. This is semantic search with a small amount of structured filtering, and it is a week or two of work on top of any modern database. Users do not call it AI. They call it search.
Bad looks like zero results because no record contains the word "plumber." Good looks like the three invoices from the plumbing vendor dated March to May, sorted by date. The filtering on dates and document type does as much work as the model.
Forms that fill themselves. Upload the receipt, photograph the license, forward the email, and the fields appear. Every field a user has to type that already exists somewhere is now a point of friction they blame on you.
The detail that separates good from bad is honesty about confidence. Highlight the fields the app is unsure of. Never silently fill a total or a date with a guess.
Summaries where there used to be lists. A thread of thirty messages, a month of transactions, a week of activity. Users expect a two-line summary at the top and the list underneath. Done well, this is the feature that gets screenshotted.
Done badly, it gets screenshotted too. A summary that invents a number is worse than no summary. Generate it from the records on screen, and let a tap on any claim jump to the source item.
Who it helps and who it hurts
It helps apps with a lot of user-generated or user-uploaded content, because that content becomes searchable and summarizable without anyone tagging it. It helps small teams, because these three features are now bought, not researched.
It hurts apps whose moat was a tidy taxonomy and a lot of manual data entry, because the taxonomy matters less when search understands sentences. It also hurts anyone who shipped a chatbot as their AI strategy and stopped there; the chatbot is now the thing users scroll past.
Picture two expense apps. One spent years refining forty categories. The other lets users photograph a receipt and type "client lunches in Denver." The second team has fewer engineers and a better product.
One prediction
By mid-2027, app-store reviews will routinely mention search quality, and apps whose search cannot handle a natural sentence will lose half a star for it. Search has become the feature users test first, and they test it by typing like a person.
The reason is habit transfer. Once people search their email and photos in plain language every day, a keyword-only box anywhere else reads as broken, not basic.
What to do this quarter
Pick one workflow where your users type the same thing twice. Usually it is data that exists in a document, a photo or a previous record. Remove the second time. That one change teaches your team the plumbing (extraction, confidence, a review step for low confidence) that every later AI feature reuses.
A sensible sequence, typically four to six weeks for a small team:
- Extract fields from the source document with a hosted model.
- Score each field's confidence and set a threshold.
- Prefill fields above it; flag fields below it for the user to confirm.
- Log every correction. That log is your accuracy measure.
Then look at your search logs. Count the queries with more than four words. If the number is growing, your users are already typing sentences and getting nothing. Fixing that is the highest-return AI work most apps have available, and it does not involve a chat window.
While you are in the logs, also count zero-result queries and queries followed by a second search within a minute. Both are users telling you search failed.
If you want a second opinion on which of your workflows to start with, ask us. The answer is usually obvious within a thirty-minute call.
Frequently asked questions
Should we remove the chatbot we already shipped?
Look at its usage. If under five percent of active users open it in a month, fold its best answers into search and help content, and reclaim the screen space.
Do these features need our own models?
No. Hosted models from the major providers, with your data retrieved at query time, cover all three. Train your own only when you have a measured accuracy problem that retrieval cannot fix.



