Parsing order confirmation emails: selectors, regex, and AI
Extract order numbers, totals, and tracking links from shop confirmation emails using CSS selectors, regex, or AI, and test everything on saved .eml files.
Every shop renders the same facts differently
An order confirmation carries a handful of facts: an order number, a total, often a carrier and a tracking link. Those facts appear in every shop's email, and no two shops put them in the same place.
A big marketplace sends nested-table HTML generated by a system from 2012. A Shopify store sends whatever the theme author decided. A small shop sends plain text written by the owner. Labels shift with language and whim: Order #, Bestellnummer, Your reference. Totals arrive as €42.90, 42,90 EUR, or EUR 42.90, with shipping sometimes folded in and sometimes on its own line.
So there is no single parser for order emails. There are three extraction tools with different economics, and a working setup usually combines them: CSS selectors for shops you buy from often, regex for clean text, and AI for the long tail.
Pick the body you parse
Most confirmations are multipart: a text/plain body and a text/html body that are supposed to say the same thing. They often do not. The text part is frequently a stub (view this email in your browser) or missing the tracking link entirely, while the HTML part is the one the template engine actually filled in.
The practical rule: extract from HTML when the shop sends real HTML, and fall back to text when the text part is complete. Decide once per sender and write the decision down. Flip-flopping between parts is where extraction bugs breed.
When you build a test corpus later, keep both parts of each message. A shop that redesigns its HTML sometimes leaves the text part untouched for years, and switching that sender from selector extraction to regex extraction becomes a one line change instead of a rewrite.
CSS selectors for the shops you order from often
Template-rendered HTML has stable structure, which makes CSS selectors the sharpest tool for a known sender. The total sits in the same table cell in every email that shop will ever send you, until a redesign, and redesigns are rare.
Real shop HTML rarely offers helpful class names, so selectors lean on structure and attributes. Attribute matching is the quiet workhorse: a carrier is easiest to catch by its tracking URL, not by the link text, which gets translated.
Selectors also fail loudly, which is a feature. When the shop redesigns, you get an empty field instead of a wrong value, and empty is easy to alert on.
# selector map for confirmations from shop.example
order_number: "td.order-meta strong"
total: "table.totals tr:last-child td.amount"
tracking_url: "a[href*='/track/']"
# extracted output
{
"order_number": "ABC123",
"total": "€42.90",
"tracking_url": "https://ship.example/track/JD2026091200418"
}Regex for text bodies and labeled values
Plain-text bodies and clearly labeled values are regex territory. The craft is anchoring: match the label and capture what follows, instead of matching anything that looks like a number. Bare digit patterns will eventually capture a street address or a customer id, and they will do it silently.
Some values have shapes strong enough to match without labels. UPS tracking numbers start with 1Z and run 18 characters. Currency amounts carry a marker and two decimals. Order numbers do not have a shape: they are whatever the shop's database emits, so anchor those to a label every time.
Expect templates to drift, and keep patterns tolerant. Allow flexible whitespace around labels, accept both decimal separators, and match case-insensitively. A pattern pinned to exact spacing breaks the day the shop switches email tools.
order_number: Order\s*#?\s*([A-Z0-9-]{5,20})
total: Total[^0-9€]{0,16}(€\s?[0-9]+[.,][0-9]{2}|[0-9]+[.,][0-9]{2}\s?€)
ups_tracking: \b(1Z[A-HJ-NP-Z0-9]{16})\b
labeled_code: Tracking (?:number|code)[:\s]+([A-Z0-9]{8,22})Where AI extraction earns its cost
Selectors and regex assume you will sit down and write a mapping per shop. For your top three shops that is an evening well spent. For the thirty others you order from twice a year it is not, and that long tail is where AI extraction earns its cost.
The instruction is written in plain English: extract the order number, the total including currency, the carrier name, and the tracking URL if present. The same instruction works across templates, languages, and shops you have never seen before, which is exactly what deterministic rules cannot offer.
The cost side is real, though. AI evaluations are metered (on HideMy.world, 250 per month on the Lite plan and 1,000 on Pro), so run the model only on mail the deterministic layer could not handle. Known shops hit their selectors and never touch the model; the unknown remainder, a handful of emails a month, goes to AI. That ordering keeps a small plan comfortable.
Test against saved .eml files
Live testing (order something, wait two days, check) is slow and unrepeatable. Build a corpus instead: export real confirmations from your mail client as .eml files, one folder per shop, and include the awkward ones: refunds, split shipments, localized templates, the order with two tracking numbers.
Then treat extraction like code under test. Upload a .eml file to run it through the transformer and compare the resulting JSON against what you expect. When a shop redesigns, add the new email to the corpus, fix the selector, and re-run the lot. Ten minutes of corpus maintenance per quarter is the entire upkeep.
Name the files by shop and case (shop.example-refund.eml, shop.example-split-shipment.eml) and keep them in the repository next to the selector maps. Extraction rules without their corpus are the kind of asset only one person on the team can safely touch.
A production setup that holds up
Give every shop the same dedicated address, say orders@yournick.hidemy.world, and let rules admit only confirmations: subject contains order, or sender matching your known shop list. Everything else arriving there is noise you never parse.
Route the clean JSON wherever it is useful: a webhook into your own order tracker, a Notion database as a purchase log, or a Telegram message the moment a tracking link appears. The payload is versioned JSON (the version is a date string, and changes within a version are additive), so a field added upstream later will not break what you built.
Two upstream details save debugging later. SPF, DKIM, and DMARC results ride along with each email and rules can gate on them, so a phishing email dressed as an order never enters your data. And Message-ID dedupe means a shop that re-sends its confirmation produces one row, not two.