Antispam MxGuard Business Product information New

Inside the LightGBM classifier: 120 features at 3.4ms

Transcom 15 Nov 2025, 23:12

The first layer of MxGuard's AI brain is a LightGBM model — a gradient-boosted decision-tree ensemble. It scores roughly 95% of all mail decisively, on its own, without calling any external service.

Speed and accuracy. Median scoring latency is 3.4 milliseconds and the model achieves an AUC of 0.988 on our test corpus. It runs in-process on the milter, so there is no network round-trip for the common case.

Around 120 features across six buckets: - Envelope — IP reputation, ASN, reverse-DNS sanity, HELO conformance, MAIL FROM domain age, SPF result. - Headers — From/Reply-To divergence, display-name impersonation, Message-ID structure, Received-chain length, DKIM result, DMARC alignment. - Subject & body — RFC 2047-decoded subject, urgency markers, gibberish measures, money/wire/invoice language, structural ratios. - URLs — link count, domain age, TLD distribution, URIBL hits, shorteners, IP-literal URLs, mixed-script hostnames. - Attachments — MIME mismatch, archive depth, macro presence, suspicious extensions, ClamAV verdict, encryption flags. - Reputation — a 7-day rolling ham/spam ratio per sender domain (Redis-cached) plus recipient interaction history.

Because the classifier is retrained continuously on a live corpus, new attack patterns are reflected in scoring within hours of appearing in the wild.

0 comments

Sign in with your TDesk account to comment.

← Back to all posts