Showing posts with label spam. Show all posts
Showing posts with label spam. Show all posts

27 September 2022

Haiku spam and double digit sigma events

Haiku Schmaiku

Howdy Ma'am,
Just spam, I am.
Five syllables short

-- Bloggerel Doggerel blog, 2007

Verse is courtesy of The Climateer, who doesn't write about climate too often, thankfully! He has a great blog description which is perpetually relevant: "In war, everything not censored is a lie."

The Climateer DOES write about investment bankers who blame statistics for their poor trading decisions... or possibly, outright deceptive practices. There was a lot of that going on in 2008. I finally hoisted some posts about double-digit standard deviations from my bookmarks and read them.

25 Sigma Event Very Unlucky

From Climateer Investor and others along the way, it seems like a 25 sigma event is impossible. "How unlucky is 25 sigma?" (2011): When Goldman Sachs was Really, Really Unlucky
"One of the more memorable moments of last summer’s credit crunch came when the CFO of Goldman Sachs, David Viniar, announced in August that Goldman’s flagship GEO hedge fund had lost 27% of its value since the start of the year."

As Mr. Viniar explained, “We were seeing things that were 25-standard deviation moves, several days in a row.” 

25 October 2013

Account hijackers

If a message originates from a familiar name or email address, its likelihood of making it through spam filters is greater.

Google described their efforts to minimize harm to users due to email account hijacking:
"Our security team...saw a trend of spammers hijacking legitimate accounts to send their messages. [We developed] a system that uses 120+ signals to...detect whether a log-in is legitimate, beyond just a password."
Less than 1% of spam emails make it into a Gmail inbox.

chart Google Gmail accounts compromised since 2010 decreased to nearly zero
Legitimate Gmail accounts blocked for sending spam versus time

The number of compromised accounts decreased by 99.7% since 2011. That's impressive, for a sustained reduction! How does Google avoid false positives? I am so curious about the specific details of their filtering rules!

The blog post was written in March 2013. It is remarkable that the same methods continue to be effective, as Gmail spam-attackers would perceive this as a new challenge to be overcome.

120 Signals


I suspect that Google's methods are analogous to those used by the U.S. Department of Health & Human Services' Centers for Medicare & Medicaid Services (CMS) in detecting medically unlikely edits (MUEs). MUEs can be accidental, due to claim coding or data entry errors. MUEs can also be deliberate, when there is fraudulent intent, e.g. by filing for more services, or for more expensive services. Regardless of intent, MUE identification reduces paid claims error rates.

How will the Affordable Care Act impact existing processes for detecting MUEs, and for setting benchmarks? CMS does not disclose its MUE criteria for the same reasons that Google will not reveal details about their 120 signals.

Continuous improvement is a part of life, for email-spam account hijackers, Google and the fraud detection team at the Centers for Medicare and Medicaid Services.

I wrote a post about health care, with a much more Ellie-centric theme, a few years ago. That was when I worked as statistician for ACCCHS, Arizona's state-administered Medicaid/Medicare program, monitoring program performance and quality of care.