AI email triage patterns that keep costs under control
Patterns for AI email triage that stay on budget: plain English questions, deterministic rules in front of the model, and designing for false positives.
Triage is classification with a budget
Email triage is one decision made over and over: this message needs a human now, later, or never. Classifiers are good at that decision, and language models are the first classifiers that work without training data. You state the question, they answer it per email.
The catch is cost discipline. Every email you show a model costs an evaluation, and plans meter them: on HideMy.world, Lite includes 250 AI evaluations a month and Pro includes 1,000. A triage design that sends everything to the model burns its budget on newsletters.
So the patterns below are less about prompts and more about plumbing: what runs before the model, what the model is asked, and what happens when it is wrong.
Scope matters too: triage decides where a message goes, nothing else. Auto-replying, summarizing, and drafting answers are separate problems with separate failure modes, and bolting them onto the routing decision multiplies the ways a mistake can reach a customer.
Write the question in plain English, and narrowly
AI rules are written as plain English questions, and the wording is the engineering. Vague questions produce vague triage: asking whether an email is important imports every ambiguity of the word important, and the model will answer confidently and uselessly.
Good questions name observable content and a narrow positive class. They read like instructions to an assistant on their first day, someone with no idea which customers are big or which merchants are yours.
Prefer several narrow rules over one broad one. Three specific questions route to three specific destinations. One question with five or-clauses routes everything to the same place and tells you nothing when it misfires.
Too vague:
"Is this email important?"
"Does this need attention?"
Narrow and answerable:
"Is the sender reporting that a payment failed or that they cannot log in?"
"Is this a new job application for an engineering role?"
"Is this an automated receipt or invoice for a purchase we made?"Keep a deterministic rule in front of the model
Classic field rules (sender contains, subject contains, body contains) cost nothing to run and are right every time within their reach. They go first, always. The model only sees what they could not settle.
A funnel for a support address might look like this: known big customers matched by sender domain go straight to the urgent channel, receipts matched by subject go to the log, and known bulk senders are dropped. For a mailbox receiving, say, 800 emails a month, rules like these can settle most of the volume, leaving the AI rule to inspect a residue that fits comfortably inside a 250 evaluation budget.
The layering also keeps the system debuggable. When a message lands somewhere surprising, you check the deterministic rules first, and most of the time the answer is sitting there in a rule you can read.
Design for false positives
Any classifier mislabels sometimes, so decide in advance which direction of error is cheap. For urgency detection, a false positive costs one unnecessary ping while a false negative costs a missed incident. Tune toward over-alerting and let the humans grumble a little.
For discard decisions the asymmetry flips: a false positive throws away someone's real email. So never let an AI verdict alone destroy mail. Route the model's safe-to-ignore bucket to a log-only destination with retention, skim it weekly, and promote anything the model got wrong into a deterministic rule.
The principle underneath: the model proposes, the architecture disposes. Pick destinations so a wrong answer is an inconvenience, never a loss.
When a dumb substring rule beats a model
If the signal is literally a string, a substring rule wins on every axis that matters: free, instant, deterministic, and explainable in one sentence. Receipts say receipt. Your monitoring tool puts ALERT in the subject. Your biggest customer writes from their own domain. None of that needs inference.
The model earns its keep where wording varies: a complaint phrased a hundred ways, my card got declined versus payment did not go through versus the same sentence in German. Matching intent across phrasings is exactly what the substring rule cannot do and the model can.
A sound habit: write the substring rule first and watch it for a while. Promote the case to an AI rule only when real emails demonstrably slip past the strings. Plenty of triage never needs the upgrade.
A quick inventory of your last hundred emails usually reveals the split. Most mailboxes are dominated by a few dozen repeat senders that strings handle completely, with a thin stream of genuinely novel mail on top. Budget the model for that thin stream and nothing else.
A worked setup on a small plan
Here is a concrete arrangement for a support address like help@yournick.hidemy.world, designed to fit Lite's 250 monthly evaluations.
The deterministic rules absorb the repetitive majority, the model reads only the ambiguous residue, and the budget holds. If volume grows, the first lever is more substring rules, not a bigger plan.
Notice what the AI rule delivers into: a channel humans watch, never an action that deletes or auto-replies. Verdicts route; they do not destroy.
- Rule 1, no AI: sender contains bigcustomer.example, deliver to Slack #support-urgent
- Rule 2, no AI: subject contains receipt or invoice, deliver to the Notion purchases database
- Rule 3, no AI: sender contains no-reply and subject contains digest, keep in the log only
- AI rule on the remainder: is the sender reporting a failed payment, an outage, or being unable to log in? If yes, send to Telegram immediately
- Everything else: forward to the verified team inbox for normal handling
Watch it, then tighten it
Run the whole arrangement in log-only mode for the first week or two: deliver to destinations nobody is paged by, and read the results like a reviewer. You are grading the question wording, not the model.
After that, review monthly. Each recurring pattern the AI catches is a candidate substring rule, and each promotion refunds evaluations for the genuinely ambiguous cases. Keep the question text in your notes under version control, because you will tune it and forget why.
Triage done this way turns into quiet infrastructure: a handful of readable rules, one carefully worded question, and a budget that stays flat while the mailbox grows.