How MxGuard is trained: honeypots, 30 years of ISP ham and public corpora
A spam filter is only as good as the mail it learns from. MxGuard's classifier is trained on three complementary sources.
- A global honeypot platform. Real spam is captured continuously from six regions — Brazil, Canada, India, Japan, Sydney and the UK — with every received message shipped back to a central corpus. This gives the model a constant, worldwide view of live attacks.
- Thirty years of ISP ham. Transcom has operated as an independent UK ISP since 1993, and its own legitimate traffic provides a diverse, realistic distribution of genuine mail. This is what keeps false positives low — the model has seen an enormous amount of ordinary, legitimate email.
- Public reference corpora. The Enron, SpamAssassin and TREC datasets provide a reproducible baseline for benchmarking.
The corpus is updated continuously, so new attack patterns are typically reflected in scoring within hours of them first appearing in the wild — a cadence a rule-based or Bayesian filter simply cannot match.

0 comments
Sign in with your TDesk account to comment.