October 1, 2026 · StartupQuickstart
AI inbox triage: where it pays off and where it doesn't
Classification, drafting and auto-sending are three different jobs with different payoffs and risks. What each costs, how to measure it, and why auto-send rarely earns its place.
AI inbox tools pitch the same promise: connect your email, and the assistant sorts, summarizes and answers it for you. Some of that promise holds up well. Some of it is a liability waiting for the wrong email. The difference comes down to which of three jobs you ask the AI to do: classify, draft, or send. They have very different payoffs and very different risks, and most disappointing projects come from treating them as one feature.
The problem is real
Microsoft's telemetry from Microsoft 365 found that the average worker receives 117 emails a day, most of them skimmed in under 60 seconds, alongside about 153 Teams messages. For the most heavily interrupted fifth of users, the same report found a meeting, email or chat ping roughly every two minutes during core hours. That is a lot of sorting, and sorting is exactly the work machines are good at.
There is a wellbeing cost too. In a two-week experiment at the University of British Columbia, 124 adults reported lower stress during the week they checked email only three times a day than during the week they checked freely, and most found the restriction hard to keep. Good triage helps here: if the urgent items are flagged reliably, people can stop refreshing the inbox to find them.
Job one: classification, where it pays off
Classification means the model reads each incoming email and assigns a label: new lead, existing customer support, invoice or billing, vendor, scheduling, newsletter, spam. Optionally it adds an urgency flag and a one-line summary. A rule then moves the email to a folder, applies a label, or notifies the right person.
This is the best place to start, for three reasons.
- It is cheap. A classification call is a short prompt and a few words of output, so it runs well on the smallest models. At Claude Haiku 5.5's price of $0.10 per million input tokens and $0.50 per million output (checked October 2026), an email of about 2,000 tokens with a 100-token answer costs about $0.00025. A shared inbox receiving 100 emails a day, 3,000 a month, costs well under a dollar a month to classify.
- Mistakes are recoverable. A mislabeled newsletter costs nothing. A lead filed under "vendor" is a problem, but it is still sitting in the inbox, unchanged, and a daily review of the lead folder catches it.
- It is easy to measure. Label 100 past emails by hand, run the classifier, and count the mismatches.
Anthropic's ticket routing guide is a good blueprint even though it is written for support tickets. It notes that an LLM can classify with just a few dozen labeled examples, that accuracy depends heavily on how well-defined your categories are, and it suggests targets like 95% accuracy across 100 test cases. It also lists the edge cases that trip classifiers up, all of which show up in real inboxes: implicit requests ("I've been waiting two weeks" is really an order status question), emotion overshadowing intent, and one email containing several requests.
Our practical advice: keep the category list short (six to ten labels), write a one-sentence definition for each, include an "unsure" label, and route anything unsure to a human. A classifier that admits uncertainty is more useful than one that guesses confidently.
Job two: drafting, where it pays off sometimes
Drafting means the model writes a suggested reply that sits in your drafts folder for a person to edit and send. The value depends almost entirely on how repetitive your email is.
It pays off when a large share of incoming email asks the same dozen questions: availability, pricing ranges, how to reschedule, where an invoice is, what to bring. The drafts can pull from a short, maintained set of approved answers, and the person's job shrinks from writing to checking.
It pays off poorly when email is mostly relationship work: a negotiation, a complaint from a key account, a delicate HR matter. Drafts there tend to be generic, and fixing a generic draft can take longer than writing a good reply from scratch. If your team deletes most drafts and starts over, the feature is costing time, not saving it.
Measure it directly. For two weeks, have the team note for each AI draft whether they sent it as is, edited it lightly, or rewrote it. If fewer than half are sent as is or lightly edited, narrow the drafting to the categories where it works and turn it off for the rest.
Job three: auto-sending, where it usually does not pay
Auto-send means the AI replies without a person reviewing the message. This is where the risk changes shape, because the model is now acting on your behalf with your name on the message.
The OWASP Top 10 for LLM applications calls this risk excessive agency: giving an AI system more functionality, permissions or autonomy than its task needs. Its own example is an email assistant. A tool meant only to summarize mail is given a plugin that can also send, and an attacker's crafted email uses prompt injection to make the assistant search the inbox for sensitive information and forward it. OWASP's recommended fixes are a read-only mail integration and requiring the user to review and send every drafted email.
Any email the assistant reads is untrusted input written by a stranger. That is fine when the output is a label or a draft that a person reviews. It is dangerous when the output is an action.
A technical detail worth knowing: in Google's Gmail API, the scope that lets an app manage drafts is described as "Manage drafts and send emails", and there is no draft-only scope in Google's list. If you build or buy a tool that should only draft, the "never send" rule has to be enforced by the software itself, not by the permission screen. Ask vendors how they enforce it, and prefer read-only access plus labels when classification is all you need.
There are narrow cases where automatic replies are reasonable: an acknowledgment that says "we received your request and will reply within one business day," sent from a fixed template that the AI does not write. That is an autoresponder with smart routing, not an AI composing messages, and it is the safe version of the idea.
Where inbox AI does not pay off at all
- Low volume. If one person gets 25 emails a day and handles them in half an hour, the setup and review time will not come back.
- Regulated or sensitive mail. Legal, medical, HR and financial correspondence carries obligations about who sees what. Sending it through a third-party model needs a privacy and contract review first, and some of it should stay out entirely.
- Undefined process. If your team cannot say who handles a billing question today, an AI cannot route it either. Write the routing rules for people first.
A rollout that works
- Baseline. For one week, track inbound volume by category and how long the inbox takes each day.
- Label offline. Hand-label 100 to 200 recent emails using your categories.
- Classify in shadow mode. Run the classifier for a week without moving anything. Compare its labels to what people did.
- Turn on routing for the categories with high accuracy, with the unsure label going to a person.
- Add drafting for one or two repetitive categories, with read access and draft creation only, and track sent-as-is rates.
- Review monthly. Misroutes, draft edit rates, and time spent in the inbox against the baseline.
Most teams find that classification alone gives them back a meaningful slice of the day, and that drafting earns its place in a few categories. Very few find that auto-send is worth the risk. If you want help scoping which of your inboxes fits which job, that is a short conversation, and the baseline week is a good thing to have done before it.
Sources
- Breaking down the infinite workday, Microsoft WorkLab (Work Trend Index special report, June 2025)
- Checking your email less can help reduce stress, University of British Columbia Department of Psychology
- Ticket routing guide, Anthropic (Claude API docs)
- Pricing, Anthropic (Claude API docs)
- LLM06:2025 Excessive Agency, OWASP Gen AI Security Project
- Choose Gmail API scopes, Google for Developers
Want systems like this built for you?
We build and run data pipelines, websites, and AI automation for startups.
