How Sift Works ebook 2025.pdf
How Sift Works
Contents
- Introduction: The Sift Technology Stack 3
- Enterprise-grade data infrastructure 4
- Massive quantities of high-quality data 5
- Sophisticated data preparation 7
- An ensemble of machine learning models 10
- A diverse ML model stack 11
- The unique qualities of Sift’s machine learning 14
- Telling a story with data 16
- Increasing efficiency through automation 16
- Summary 17
Introduction: The Sift Technology Stack
In 2011, Sift disrupted the fraud prevention industry with a first-of-its-kind machine learning approach that accurately predicted fraud and defended against online abuse in real time.
Today, our global data network, custom machine learning models, automation technologies, and comprehensive reporting fuel the business growth of tens of thousands of websites, allowing them to prevent fraud, streamline operations, and drive revenue growth.
From a technology perspective, detecting fraud is extremely difficult – like finding a needle in a haystack. Additionally, fraudsters are constantly evolving and adapting their techniques. So how do we do what we do effectively? We’ve designed our technology stack to meet these challenges head-on.
Sift receives billions of events per month and analyzes vast streams of data in real time. Our platform hosts a powerful machine learning engine that allows us to detect and prevent fraud by analyzing nuanced combinations of signals buried in historic and real-time data.
Our approach combines speed, scale, and sophistication to deliver a unique, adaptive solution that allows our customers to accurately distinguish between the users they can trust and those they can’t. Savvy businesses use this knowledge to focus on improving the experiences for trusted users, while keeping fraudsters at bay.
Let’s take a look into the powerful technology at the heart of the Sift engine.
Enterprise-grade data infrastructure
Sift offers a secure, reliable, and scalable infrastructure that opens up multiple integration points to ensure the successful capture of critical data from any source that you use.
Our solutions specialists will ensure that you make the most of Sift. As trusted partners, they will guide you as you integrate your desktop and mobile experience using our Javascript snippet and SDKs, and connect your backend systems using our REST APIs. Our easy to use integration guides and support center articles ensure that you can also find your own answers along the way. Many organizations are able to get started with Sift with just a few hours of engineering time.
Data infrastructure drivers
Sift has three data infrastructure drivers: reliability, scalability, and security. Let’s take a look at each
Reliability
Sift supports a fast-growing portfolio of enterprise-level and international clients. We strive to provide a resilient and highly available service to our customers across the globe. Hosted on Google's Cloud Platform (GCP), Sift employs a suite of fault-tolerant features aimed at eradicating single points of failure, including a deployment across multiple Availability Zones, real-time replication services, and nightly backups. These are managed through 24x7 continuous monitoring and supported through regularly updated and tested runbooks, business-continuity and disaster-recovery procedures, and incident-response plans. You can always check our system status by visiting our public status page.
Scalability
Security
We’re serious about protecting data. Sift maintains compliance with the SOC 2 framework. We employ strict access control, two-factor authentication, and encryption of data in-flight and at-rest to ensure that customer data is protected. Our annual SOC 2 Type 2 assessment tests our security, availability, and confidentiality processes and controls against the SOC 2 security framework, ensuring independent, third-party assurance that we are taking steps to protect our systems and our customers’ data.
Massive quantities of high-quality data
Machine learning requires access to massive quantities of high-quality, relevant data in order to be truly effective at detecting and preventing fraud. The quality of the results you see are directly in line with the quality of data sent to Sift.
We are constantly pushing the envelope to discover hidden fraud patterns before they can damage your business. To accurately predict fraudulent behavior in real time, we sift through a variety of data types and formats that together create a holistic picture of emerging fraud threats.
Raw data
In the world of machine learning, the most accurate analyses come from high-quality data. We centralize an array of datasets gathered via SDK, Javascript snippet, and API – augmented with a variety of relevant third-party data – so that our customers get the most comprehensive view of the world. By integrating with Sift, the burden of data collection and management is taken off of your plate. The data that we leverage can be broken down into the following categories:
USER IDENTITY
Attributes that are associated with the identity of a user
Examples: Name, email address, phone numberBEHAVIORAL PATTERNS
Preferences and patterns associated with the event
Examples: Browsing patterns, URL information, screen height/widthLOCATIONAL DATA
Location attributes associated with the event
Examples: Fine and coarse location, GPS coordinates, shipping address, billing addressDEVICE & NETWORK DATA
Properties of the device and network connection associated with the device
Examples: IP Information, Network ID, carrier network, device manufacturer & modelTRANSACTIONAL DATA
Order details and order history associated with the event
Examples: Order value, order velocity, payment instrumentsDECISIONS
Feedback on Sift's fraud and risk predictionsCUSTOM DATA
A variety of relevant third-party datasets
Examples: IP information, BIN information, currency rates and conversions
Knowledge derived from data
Organizing data for the prediction tasks helps to establish connections between users, devices, locations, and other attributes. When we look at data, we don’t just review the last event, but rather analyze the entirety of behavior and actions over a span of time. This broad-scale perspective allows us to learn new user behavior relative to other users for each customer, and scale those learnings across the entire network. This time series nature of modeling affords us a great deal of flexibility and utility when fed into the machine learning system.
Some of the knowledge examples that provide valuable insights into the behaviors of fraudsters include:
- TIME SERIES DATA
Sophisticated data preparation
Suspicious behavior is often buried within streams of data. Sift has built a sophisticated system that combs through the world’s largest database of fraud-related data collected from our global network in real time, and maps it against meaning, relationship structure, and relevancy – often in company-specific ways. This evidence informs our machine learning models that recognize new patterns and update in real time. Sift pioneered this approach and has proven its success time and again with businesses of all sizes and industries across the world.
Data normalization
Fraudsters are constantly hunting for new methods to get around existing system controls and rules. That means that as fraudsters adapt their tactics, businesses can be vulnerable to new types of fraud attacks. Starting with data normalization, Sift spots the little details that other approaches miss. Here are two examples of data normalization techniques that we frequently use:
As an example, a customer might use a rule or a blacklist to block a fraudulent email address, e.g. johndoe123@gmail.com. In response, a fraudster will often create a similar-looking email address, e.g. johndoe124@gmail.com, to circumvent the controls enforced by your system. A similar technique is common with physical addresses:
Ralph Wiggum
123, Smith Ln
San Francisco, CAjohndoe123@gmail.com
johndoe124@gmail.com
johndoe_123@gmail.comRalph W
Smith Lane, #123
San Francisco, CaliforniaRAW DATA
Ralph C Wiggum
#123, Smith Ln
San Francisco, CA
Variations of the same physical shipping addresses used by fraudsters
Therefore, when we spot minor variations to a known fraudulent email address, we are able to accurately match and flag similar email addresses.
Currency conversion
Currency conversion is another critical data point. If a buyer on your site spends money in a currency different from the majority of your users, we will properly adjust the order amount on our end before comparing it with other orders. Since currency conversion rates fluctuate continually, we base this comparison on conversion rates around the time of the sale, not on a single fixed currency conversion chart.
Feature engineering
A key challenge in building an effective machine learning system that accurately detects a variety of fraud vectors is feature extraction – deriving the most useful signals from all kinds of raw data sent to Sift. Feature engineering transforms raw data into structured, machine-processable formats that can be understood by a machine learning algorithm. Why is this important? It allows us to set up building blocks that are powerful indicators of fraud. For example, a count of the number of vowels per email address when applied to a machine learning model could be used as a strong fraud signal.
Large-scale learning
Fraud isn’t static, and new patterns emerge daily. Building an effective set of features that will uncover fraud indicators capable of detecting and blocking tricky behavior requires deep knowledge of the industry, our customers, their end users, and fraudsters. With feature engineering, experience is everything. Sift has a library of over 16,000 features that we use to uncover fraud patterns across many industries and time zones. Analysis of false positives and false negatives identified by our customers further contributes to and greatly improves our detecting capabilities, and those findings are used by machine learning models across the entire network.
An ensemble of machine learning models
Each business is unique. We make every effort to customize our approach to catch the fraud that is specific to your needs. Say you’re a shoe company – we might recognize that buying size 15 shoes at a particular time of day is associated with fraudulent activity. Sift’s flexible machine learning models can detect such patterns of normal and abnormal behavior that are very unique to your business model, industry, and audience.
At Sift, instead of relying on a single machine learning model, we use an ensemble of several predictive models. Some models are trained with a general understanding of fraud patterns across our network of customers, and others are tuned to your organization’s data specifically. This ensemble of models allows Sift to accurately score a transaction or a session while taking a holistic approach when analyzing risk.
The unique qualities of Sift’s machine learning
Real-time learning is enabled through learnings from across the global network of websites and applications that use Sift, and from custom models tailored to customer needs. We share knowledge about fraudulent behavior so that together we become smarter and better equipped to keep fraudsters at bay. The information is pushed in real time to update our network models and is propagated to all of our customers in just a few milliseconds.
Increasing efficiency through automation
- Workflows
Build and manage your business logic within Sift. Take automatic action on events based on a set of customizable criteria (e.g., risk score > 80, first-time user, country=Japan, etc). - Review Queues
This feature is the most efficient way for your analysts to manually review orders and events. - Sift Connect
Initiate actions in third-party services such as dispute resolution and verification from within Sift.
Summary
Building a highly accurate system for preventing online fraud is a complex endeavor that requires constant monitoring, tuning, and engineering resources. You must be able to ingest large volumes of high-quality data, use that data in various real-time machine learning models and algorithms, manage automated business logic and decisions, and enable your review teams to investigate and act with speed and accuracy. Here at Sift, we’re always thinking about how to empower our customers and have taken into account these many challenges and more. Businesses leverage Sift to prevent multiple types of fraud, while creating outstanding customer experiences.