RevSure Predictive AI Models Overview

Prev Next

Introduction

As B2B go-to-market strategies grow increasingly complex, the need for data-driven decision-making has never been more critical. RevSure’s modeling framework blends heuristic logic and machine learning to deliver predictive insights across the entire revenue funnel, from account prioritization and lead scoring to pipeline forecasting and campaign effectiveness.

This document outlines the overall  modeling approaches that power RevSure’s AI-driven GTM engine. It explains how different models are used based on data availability, describes the logic behind each predictive module, and details how outputs are scored, categorized, and used to drive marketing, sales, and RevOps efficiency. Whether navigating low-data environments or optimizing a mature funnel, RevSure ensures you have actionable intelligence to drive conversions, improve forecasting accuracy, and unlock incremental growth.

Note on Modeling Approaches:

Wherever applicable, RevSure uses both heuristic and machine learning approaches. Machine learning is the default when sufficient historical data (typically 2–3 years) is available. In scenarios with limited history, heuristic logic ensures reliable predictions. Heuristic in our terminology means statistical and advanced statistical approaches informed by business inputs.

In many cases, we also reconcile heuristic and machine learning approaches to align top-down macro GTM motion trends and conversion variations with bottom-up ML-based predictions.

Further, RevSure is constantly improving its modeling approaches, feature engineering, training and inferencing methods. At any time this document only reflects the overall approach.  All our models follow robust data science and machine learning best practices from training dataset preparation, data imputations, feature engineering, feature selection, hyperparameter tuning, hold-out  validations and scoring.

AI-based Account Prioritization Model

The Account Prioritization Model is an AI-driven solution designed to enhance sales and marketing efficiency by identifying and focusing on accounts with the highest likelihood of conversion.

It evaluates multiple data points to assign a propensity score to each account, enabling teams to prioritize efforts effectively and focus on prospects likely to advance through the sales funnel and drive revenue.

How It Works

Machine Learning based on Historical Data

The model incorporates both Base Fit Propensity and 3-Month Propensity to prioritize accounts more effectively. By analyzing historical behavior, engagement patterns, and pipeline dynamics, it predicts a base propensity score, based on ICP characteristics like industry, region, and company size as well as the likelihood of conversion within the next 3 months.

1. Base Fit Propensity

The Base Fit Propensity evaluates the long term and ICP fit of an account based on a Machine Learning model that is trained on historical conversion behaviour. It measures the inherent compatibility of an account based on demographic and firmographic attributes, such as:

  • Industry

  • Company Size

  • Region

  • Firmographics…

  • Etc.

Approaches:

  1. Heuristic Method (for low data scenarios where at least 2 years of history is not available) :

    • A probability based approach that uses historical data and predefined metrics to assess the likelihood of conversion.

  2. Machine Learning Model (for all scenarios where we have at least 2 years of history)

    • A more sophisticated approach leveraging machine learning to predict Base Fit Propensity.

    • The ML model analyzes historical patterns using algorithms like XGBoost Classifier.

The resulting scores indicate the likelihood of an account converting to pipeline or bookings.

2. 3-Month Propensity

The 3-Month Propensity predicts the short-term likelihood of conversion and is entirely machine learning-based. This propensity score incorporates:

Account Characteristics:

  • Region

  • Industry

  • Company Size

  • Existing Customer Status

  • Etc.

Funnel Stage, Pipeline, and Opportunity Metrics:

  • Volumes in stages

  • Open pipeline counts

  • Etc,

Engagement, Momentum and Activity Metrics:

  • Count activity types by lead (e.g., website visits, campaign interactions,  whitepaper downloads, etc.)

  • Journey features : sequence of interactions by campaign types, channels, etc.

  • Account-level lead counts

  • Count of activity types by opportunity

  • Volumes of contacts and engagement as of a specific date

  • Etc.,

Temporal Features:

Date-based features such as day, month, quarter, and year, etc.  are extracted from account activity timestamps.

Algorithm:

Due to its efficiency and robustness, the XGBoost Classifier is used as the core algorithm for generating these short-term predictions. RevSure also uses ensemble approaches as needed if it helps significantly improve classification predictions (F1 score, etc.)

Retraining Frequency: Once every quarter, or earlier if there are changes in data configurations.
Scoring Frequency: Daily

3. Scoring and Prioritization

The model results in the Base Fit Propensity and 3-Month Propensity as a comprehensive score for each account. These scores are categorized as:

  • Pipeline Fit Propensity: Indicates the account’s likelihood of contributing to the pipeline in the long term.

  • Booking Fit Propensity: Indicates the account’s likelihood of contributing to bookings in the long term.

  • 3-Month Pipeline Propensity: Indicates the likelihood of pipeline generation within three months.

  • 3-Month Booking Propensity: Indicates the likelihood of converting to bookings within three months.

Each score is further assigned into buckets such as High, Medium, Low, or Deprioritize, based on the model’s confidence and the account’s potential.

Structured Output and Insights

The output of the model is a structured, ranked list of accounts with detailed scores and categorizations. For example:

  • Fit Propensity (Pipeline & Booking): A measure of long-term fit.

  • 3-Month Propensity (Pipeline & Booking): Short-term likelihood of success.

Actionable Steps:

  • Focus on high-priority accounts for immediate follow-up.

  • Nurture medium-priority accounts.

  • Deprioritize low-potential accounts to optimize resource allocation.

Lead Propensity

The Lead Propensity Model estimates the likelihood of a lead moving to the next stage in the funnel. By analyzing past behaviors and funnel transitions, we assign each lead a probability score that reflects its potential for conversion.

How Scores Are Assigned: Lead Propensity could be generated using heuristic or machine learning approaches to assign scores. By default, we use the machine learning model, provided there is at least 2–3 years of historical data available. If sufficient history isn’t available, we fall back on the heuristic method to ensure reliable estimates.

In the heuristic version, historical conversion rates and average opportunity sizes are computed for different lead segments (e.g., by company size, industry, and various persona). These values are then used to calculate the projected value by multiplying the number of open leads with historical conversion and opportunity size estimates (more about this is mentioned in the pipeline projection section).

The machine learning version trains a classification model on enriched lead data to estimate conversion probabilities. The model learns from patterns of previous lead behaviors, such as which leads converted in the past, how quickly, and with what factors being impactful. The model outputs a probability that reflects the likelihood of the lead progressing to the next stage. These scores are then computed across multiple quarters  (CQ, NQ, NQ+1, NQ+2), representing the likelihood of conversion within each specific quarter.

Data Inputs (Overview)

These data inputs are descriptive. 100’s of features go into the ML model

  • Lead attributes: source, region, company size, industry, etc.

  • Activities and campaign interactions (multiple features)

  • Journey features : sequence of interactions by campaign types, channels, etc.

  • Funnel stage  timestamps: lead creation date, MQL date, SQL date

  • Derived metrics: stage duration, time to quarter-end, velocity of movement

  • Etc.

Feature Selection: Feature Importance, Chi-squared test, ANOVA F-test

Retraining Frequency: Once every quarter during Projection Engine training, or earlier if there are changes in data configurations.


Scoring Frequency: Daily

Interpretation:
A higher score indicates a greater chance of the lead moving to the next stage. Teams should prioritize outreach for leads with high scores and review or nurture those with lower scores.

Opportunity Propensity

The Opportunity Propensity Model predicts how likely an opportunity is to close within each of the following four quarters. It considers the opportunity's lifecycle and characteristics to estimate conversion likelihood over time.

How Scores Are Assigned: The scoring methodology for Opportunity Propensity closely mirrors that of Lead Propensity but is enriched by a broader and deeper set of data. While both models learn from past conversions. Opportunity Propensity incorporates more granular and diverse features, such as opportunity type, forecast category, and detailed stage progression, to generate more context-aware predictions. Each opportunity is scored for the following four quarters (CQ, NQ, NQ+1, NQ+2), representing the probability of its closing in each respective quarter.

Data Inputs:

  • Opportunity creation date, close date, forecast category, opportunity type, CRM-provided probability, focus quarter features, i.e., how far is the assigned close date quarter.

  • Stage data: current and past stage timestamps, stage sequence

  • Journey features : sequence of interactions by campaign types, channels, etc.

  • Derived metrics: days in stage, time to close, time to quarter-end

  • Etc.

Feature Selection: Feature Importance, Chi-squared test, ANOVA  F-test

Retraining Frequency: Once every quarter during Projection Engine training, or earlier if there are changes in data configurations.
Scoring Frequency: Daily

Interpretation:
Scores help assess which deals are most likely to close. Focus attention on opportunities with substantial probabilities in the current or upcoming quarters.

Opportunity Push and Pull Models

The Push and Pull models are designed to improve pipeline readiness projections by accounting for opportunity close date shifts across quarters. Traditional pipeline estimates only consider opportunities with close dates in the target quarter, ignoring the fact that opportunities often move between quarters.

  • Pipeline Pull: Opportunities originally scheduled to close in other quarters (Qb) but are likely to shift into the quarter of interest (Qa).

  • Pipeline Push: Opportunities scheduled in the current quarter (Qa) but are likely to shift out to other quarters (Qb).

By modeling these transitions, we can create a risk-adjusted pipeline that more accurately reflects what will realistically be available for closing in any given quarter.

Key Steps and Principles

  • Data Preparation

  • Track all historical close date changes of opportunities.

  • Especially consider movements where the close date shifts across quarters

  • Transition Tracking:

    • Capture transitions in the form: Close Date Transition Probability: Qb → Qa during Qc

    • Qb: original close date quarter

    • Qa: target close date quarter

    • Qc: relative quarter at evaluation time (e.g., Q+0, Q+1, Q+2, Q+3)

    How it Works

    1. Determine Qc:

  • Based on the current date and opportunity creation date, establish whether the opportunity is in Q+0, Q+1, Q+2, etc.

    2. Estimate Opportunity Movement:

  • Use the ML models to estimate whether the opportunity will move during Qc.

  • This involves two ML models:

    • Movement Prediction Model – predicts whether the close date will move at all.

    • Transition Prediction Model – if movement occurs, predicts the probability to which quarter (Qa) the close date will shift.

    3. Calculate Pull (Pipeline from Other Quarters into Qa):

  • For each opportunity:

    • Multiply the probability it moves during Qc by the probability it transitions from Qb → Qa.

  • Aggregate across all opportunities and Qbs to estimate:

    • Pipeline Volume (count of opportunities)

    • Pipeline Value (weighted by opportunity amount)

    4. Calculate Push (Pipeline from Qa to Other Quarters):

  • Similar approach, but tracking opportunities moving out of Qa into other quarters.

    5. Reconcile Movements:

  • Apply both heuristics and ML reconciliation for accuracy.

  • Outputs:

    • Probabilities of the close date movement to each of the quarters Q+1, Q+1, Q+2, etc.

Scoring Frequency: Daily

Lead Fit Model

RevSure's Lead Fit model estimates how closely a newly created lead matches the profile of leads that have historically converted into pipeline or bookings. Unlike propensity models, which rely on engagement signals accumulated over time, the Lead Fit model uses demographic and firmographic attributes available at the time of lead creation to provide an immediate assessment of lead quality. This enables sales and marketing teams to prioritize promising leads from the outset, even before meaningful engagement data becomes available.

Modeling Technique:

The model learns historical conversion patterns across key demographic and firmographic attributes, including:

  • Department

  • Seniority

  • Industry

  • Company Size

  • Geography / Region

  • Lead Source

  • Other tenant-configured lead attributes

For each attribute value, the model calculates how frequently similar leads have historically converted to the target stage. Rare attribute values are grouped together to improve statistical reliability.

When scoring a new lead, the model retrieves the historical conversion likelihood associated with each of its attributes and combines them to produce an overall Lead Fit score. If an attribute value has not been observed previously, the model falls back to the overall historical conversion rate, ensuring every lead receives a score.

Outputs:

  • Lead Fit Score: likelihood that the lead matches the profile of historically successful leads.

  • Normalized Lead Fit Score: score scaled to a consistent range for easier comparison and prioritization.

ReTraining Frequency: Quarterly

Scoring Frequency: Daily

Markov Chain based MTA

The Markov Chain–based multi-touch attribution (MTA) model assigns conversion credit across all marketing touchpoints by analyzing the probability of a customer’s journey  through different channels. Unlike simple rule-based models (first-touch, last-touch, linear), it uses probabilistic modeling to capture the impact of each channel in driving conversions.

Key Steps and Principles

  • Customer journeys are modeled as a Markov Chain, where each touchpoint is a state.

  • A transition matrix is built to represent probabilities of moving from one touchpoint to another.

  • Attribution is calculated using the removal effect: the drop in conversion probability when a channel is removed.

  • Data Preparation

  • Touchpoints: unique marketing channels (e.g., Email, Paid Search, Social).

  • Conversion Paths: sequences of touchpoints leading to conversions.

  • Transition Matrix: probabilities derived from these paths, showing how users progress toward conversion.

  • How it Works

  1. Construct the transition matrix from historically observed paths.

  2. Simulate journeys until conversion or drop-off.

  3. Estimate each touchpoint’s contribution by measuring how often it leads toward conversion.

  4. Normalize scores so contributions across channels sum to 100%.

  • Outputs

  • Shows the key customer paths and their corresponding conversion probabilities

  • Show contributions to stage by stage conversion for each channel, campaign type, etc.

Scoring Frequency: Daily

Hidden Markov Models based MTA

RevSure’s Hidden Markov Model (HMM)–based multi-touch attribution (MTA) enhances traditional probabilistic attribution by incorporating the firmographic and demographic attributes of leads and accounts — such as region, industry, company size, revenue, and persona — along with the characteristics of each campaign touch, including channel type, message theme, offer, and timing. This allows RevSure to uncover not just which channels drive funnel progression, but why specific combinations of audience attributes and campaign characteristics are most effective at different stages of the buyer journey.

Key Steps:

Data Preparation

  • Touchpoints: all marketing and sales interactions — including Email, Paid Search, Social, SDR Outreach, and Events.

  • Funnel Stages: aligned with the funnel definition at the customer (e.g., Inquiry → MQL → SAL → SQL → Opportunity → Closed Won).

  • Training Data: historical lead and opp journeys with timestamps, sequence of touches,  campaign metadata, attributes campaign details, and conversion outcomes.

  • Firmographic and Demographic Enrichment: each journey is enriched with lead and account-level attributes such as region, industry, company revenue, employee size, and persona.

  • Campaign-Level Attributes: each touchpoint includes metadata like channel type, content format, offer type, message theme, etc. allowing the model to differentiate performance across creative and tactical dimensions.

Modeling Technique:

RevSure uses a Hidden Markov Model (HMM) to learn probabilistic patterns from historical customer journeys. Unlike traditional Markov models that treat each touchpoint as a state, HMM models journeys as sequences driven by underlying patterns, enabling it to better capture dependencies across interactions.

The model estimates the likelihood of different touchpoint sequences leading to conversion, incorporating both channel transitions and contextual factors such as campaign characteristics and audience attributes.

Using these probabilities, RevSure evaluates conversion likelihood for complete journeys. Attribution is derived using a removal-based approach, where each touchpoint is excluded from the sequence and the resulting change in conversion probability is measured.

This change represents the touchpoint’s incremental contribution, and attribution weights are assigned accordingly to reflect true influence rather than positional bias.

Compared to standard Markov Chain models, this approach:

  • Captures more complex interaction patterns across touchpoints

  • Incorporates campaign- and audience-level context

  • Better reflects how combinations of touches drive conversion outcomes

  • Models the full customer journey in a single framework, capturing interactions across stages (unlike stage-wise Markov approaches)

Training Frequency: Quarterly

Scoring Frequency: Daily

Account Buying Stage

RevSure's Hidden Markov Model (HMM) based Account Buying Stage model identifies where each B2B account sits in its buying journey (e.g. Awareness, Research, Consideration, Evaluation, Decision) by learning from how engagement, pipeline, and GTM signals evolve week over week. Unlike static CRM-based stage assignments, it treats buying stage as a latent state inferred from observable sequences of marketing activity, campaign interactions, and funnel progression, and accounts for the natural dynamics of B2B buying, where accounts may advance, stall, or briefly regress across stages over time. This allows RevSure to deliver an interpretable, continuously updated stage label for every account, reflecting its actual behavioral trajectory rather than a point-in-time snapshot.

Data Preparation

Training data is built as an account × week time series, where each row represents one account in one weekly window, spanning roughly the past year of history. Features fall into three broad categories:

  • Engagement & campaign signals: total activity counts, email/form/web interactions, campaign touches by channel and type, windowed engagement trends, and high-intent interaction ratios

  • Pipeline & funnel signals: counts at each funnel and pipeline stage, time-in-stage metrics, opportunity age, and past pipeline/booking conversions

  • Account attributes: industry, segment, company size, and existing customer flag where configured

Raw daily data is rolled up into 7-day windows. Cumulative metrics use end-of-week values. Week-over-week deltas and log-transformed activity counts are derived before modeling to capture velocity and stabilize scale.

Modeling Technique:

The HMM treats the buying stage as a hidden state inferred from observable weekly feature vectors. It learns three things from historical account sequences: the typical engagement and pipeline patterns associated with each stage (emission parameters), the likelihood of an account staying, advancing, or regressing between stages week over week (transition matrix), and the distribution of starting stages (initial state probabilities).

To keep stage ordering interpretable and aligned with real buyer behavior, domain constraints are baked in before fitting. States are initialized by anchoring early stages to low-engagement accounts and the final stage to closed-won events, with middle stages interpolated in between. Parameters are estimated via Expectation-Maximization (EM), and at inference time the Viterbi algorithm decodes the most likely stage sequence for each account's recent history.

Outputs:

  • Account buying stage: predicted current stage for the account

  • Buying stage confidence_score: posterior probability of that stage assignment (0–1)

  • Stage transition history: weekly stage sequence per account

Training Frequency: Quarterly

Scoring Frequency: Weekly/Daily (depending on the deployment)

Title Standardization Model

RevSure's Title Standardization model converts diverse job titles into standardized Seniority and Department classifications, enabling consistent reporting, audience segmentation, buyer persona identification, and downstream AI models. By normalizing variations in job title terminology across different data sources, the model ensures that equivalent roles are classified consistently while supporting tenant-specific organizational taxonomies.

How the Methodology Works:

The model uses a layered classification approach. Previously standardized titles are first resolved through exact and semantic matching against an existing repository of standardized titles. New or unseen titles are then classified using a Large Language Model (LLM), which maps each title to exactly one seniority level and one department from the tenant's configured taxonomies. Newly classified titles are automatically added to the repository, improving consistency and reducing future processing.

Outputs:

  • Standardized Seniority: normalized seniority level

  • Standardized Department: normalized functional department

Training Frequency: The standardized title repository is continuously enriched as new titles are processed.

Scoring Frequency: Daily

Buyer Persona Model

RevSure's Buyer Persona model classifies individual contacts into buying roles that reflect how they typically participate in a B2B purchase decision- such as Economic Buyer, Technical Buyer, Champion, or Gatekeeper. It operates at the contact level, identifying each person's role in the buying process. Persona definitions, dimension weights, and matching strategies are fully configurable per tenant through the RevSure platform- including the ability to define alternate mappings scoped to specific industries, segments, regions, or account tiers, so buying group definitions remain accurate across diverse account portfolios.

Modeling Technique:

Classification is driven by person-level attributes- typically job title, seniority, and department combined with the tenant's persona configuration. Before scoring, the pipeline resolves which persona criteria apply to each contact. Context rules allow title, department, and seniority definitions to be overridden for specific account dimensions such as industry, segment, region, or tier. Rules are evaluated top-down; the first match wins, and the base config applies if none match. The resolved criteria then flow into dimension matching and scoring.

How the Methodology Works:

  1. Dimension matching: Each configured dimension (title, seniority, department, or custom) is scored against the persona's criteria. Four matching strategies are supported:

  • Semantic: embedding-based similarity for varied or informal text

  • Fuzzy: near-exact string matching with minor variation tolerance

  • Exact: normalized string match for standardized values

  • Rule-based: ladder-style matching for hierarchical fields like seniority

Scores are normalized to 0–1 per dimension.

  1. Persona scoring: Dimension scores are combined as a weighted average to produce a score for each persona. Weights are re-normalized over only the dimensions that have values for that contact, so missing data does not penalize the assignment.

  2. Classification: The highest-scoring persona is assigned if it clears a minimum score threshold and a minimum margin over the second-best persona. Contacts that don't meet these gates are labeled Unclassified.

  3. Business guardrails: Configurable exclusion rules ensure persona assignments remain consistent with seniority and role expectations, preventing implausible matches at the edges of the distribution.

Outputs:

  • Buyer Persona: assigned persona label (e.g. Champion, Economic Buyer)

  • Buyer Persona Score: overall match score for the assigned persona (0–1)

Demand Generation Potential

This model forecasts the additional pipeline and booking value from outside the funnel. These unseen opportunities can originate from new leads or ongoing activities. Since these deals do not yet exist in the system, the projection is made at a quarterly level rather than at the individual lead or opportunity level.

How the Methodology Works: We use heuristic and machine learning approaches to estimate unseen contributions.

1. Heuristic Approach:

  • For each day in a quarter, we calculate the proportion of pipeline or bookings that originated from leads/opportunities not present at the beginning of the quarter.

  • We apply this average on a rolling basis from the current day to the end of the quarter to project the amount of additional pipeline or bookings that can be expected.

2. Machine Learning Approach:

  • A regression model is trained using historical patterns of unseen contributions.

  • The model uses date-based features and cumulative actual pipeline/booking data to predict unseen volume/value percentages.

Features Used in the Regression Model

    • Spend and impressions across channels and campaigns

    • Active sales reps and/or BDRs

    • Macroeconomic factors like interest rates and GDP growth

    • Calendar attributes such as day-of-quarter, day-of-month, week-of-quarter, month-of-year, etc.

    • Expected time series forecasts for the quarter

Interpretation:

This model provides a forward-looking estimate of pipeline and bookings that will likely materialize later in the quarter.

Product Prediction Model

This machine learning model predicts the products likely to be associated with the lead and account when it converts to an opportunity.

The product prediction model is used downstream to then also predict the Opportunity size associated with lead and account when it converts to an opportunity.

Data We Use:

  • Lead and Account attributes: account size, industry, region,segment, existing customer (y/n)

  • Product name and product family list from the Opportunity history

Retraining Frequency: Once every quarter during Projection Engine training, or earlier if data configurations change
Scoring Frequency: Daily

Opportunity Size Prediction

This module helps estimate the likely dollar value of a deal based on what we already know about the opportunity. Not every opportunity, especially those in the early stages of the funnel, has a reliable booking amount. This model helps fill that gap.

How Scores Are Assigned:

We use two complementary approaches:

  • Heuristic Approach: Looks at similar past opportunities grouped by key attributes (like industry, company size, product) and calculates an average booking amount. This value is then assigned to open opportunities that share those attributes.

  • Machine Learning Approach: Learns from patterns across a broader set of features (e.g., lead source, persona, product history, company size, region, industry, etc.) to predict the most likely booking amount for each opportunity.

Data We Use:

  • Lead and Account attributes: type, industry, company size, lead source, country, etc

  • Historical booking values and conversion patterns

  • Product details and predicted associations from the Product Prediction model  (if available)

Feature Selection: Information Value

Retraining Frequency: Once every quarter during Projection Engine training, or earlier if data configurations change
Scoring Frequency: Daily

Interpretation:

This model ensures that each opportunity, even if incomplete, has a realistic projected dollar value. This enhances downstream projections and provides a stronger view of revenue potential.

Pipeline Projection

The Pipeline Projection module clearly shows how much pipeline and bookings are expected to convert in the upcoming quarters. It helps answer the question, "How ready is our pipeline to convert to real revenue?"

What It Does: This model shows both pipeline readiness and booking readiness. It aggregates the projected value and volume of deals expected to close for each of the next four quarters (CQ, NQ, NQ+1, and NQ+2).

How It Works:

  • It uses the conversion probabilities generated from the Lead and Opportunity Propensity models.

  • Each lead or opportunity is assigned a projected value using the Opportunity Size Prediction model.

  • Then we take the product of the value calculated and the likelihood of conversion to estimate its contribution.

  • These individual contributions are aggregated by quarter to produce overall pipeline and booking projections.

  • The projections undergo careful scaling, informed by the output of a predictive model used for macro adjustments.

  • In addition to the previously calculated contribution, the Demand Generation Potential observed for the quarter from the corresponding day is subsequently included.

Feature Selection: Feature Importance and Shapely value combination

Retraining Frequency: Once every quarter, or earlier if there are changes in data configurations.
Scoring Frequency: Daily

Interpretation:

This model gives a quarterly view of how much pipeline and bookings will likely materialize, combining projected readiness volume and value.

Macro Forecasting Methodology

The Macro Forecasting model offers a clear view into future pipeline and booking expectations by predicting values for CQ, NQ, NQ+1 & NQ+2 for key metrics: Generated Pipeline Volume, Generated Pipeline Value, Generated Booking Volume, and Generated Booking Value.

Objective

The model estimates end-of-quarter pipeline and booking potential across the next four quarters, providing visibility into upcoming demand generation capacity.

Key Steps and Principles

Data Preparation

  • Collect historical records of sales activity, SDR/BDR activities, marketing campaigns, channels, spends, and macroeconomic indicators.

  • We use daily-level datasets with at least two years of historical data.

  • Engineered features such as:

    • Active sales reps or BDRs

    • Spend and impressions across channels and campaigns

    • Open pipeline volume and value at each stage

    • Macroeconomic factors like interest rates and GDP growth

    • Calendar attributes such as day-of-quarter, day-of-month, week-of-quarter, month-of-year, etc.

    • Expected time series forecasts for the quarter

Model Training and Selection

  • Separate models are trained for each combination of metric (value/volume) and funnel stage (pipeline/booking), and for each quarter (CQ to NQ+2).

  • Machine learning regression models are trained to address specific forecasting objectives. These models automatically manage outliers and are tuned to maximize predictive performance.

Forecast Scoring and Output

  • Trained models are applied daily to the most recent data to generate forecasts

  • Forecasts show the end-of-quarter expectations from the latest date in the current quarter (CQ) for CQ, NQ, NQ+!, NQ+2.

Output Utilization in Pipeline Projection

The macro forecast model is a key component of the broader "Pipeline Projection" system. While record-level predictions capture the readiness of individual leads and opportunities, the macro forecast ensures these projections are grounded in larger patterns observed across marketing and sales activity.

  • Conversion probabilities from the Lead and Opportunity Propensity models are multiplied by predicted opportunity values (from the Opportunity Size Prediction model).

  • These record-level projections are aggregated by quarter to generate raw quarterly forecasts.

  • The macro forecast model provides scaling factors based on historical patterns, economic trends, and demand generation behavior, both at the Volume and $ Value level.

  • The aggregated projections from record-level projections are compared against the macro model forecasts to come up with a scale factor that is applied to the record-level projections.

  • Record-level projections are then adjusted by the scale factor such that the aggregated quarterly projections match the macro model forecasts.

  • This reconciliation process between the micro and macro models ensures that the total projections in pipeline readiness and booking readiness modules reflect realistic, quarter-aligned expectations.

Interpretation

  • The model’s output refines the projected value and volume of pipeline and bookings for CQ, NQ, NQ+1, and NQ+2.

Retraining and Scoring Frequency

  • Retraining: Once per quarter in sync with the projection engine; earlier retraining may occur when there are configuration or data schema changes.

  • Scoring: Daily, using the most up-to-date information.

Campaign Performance Prediction Engine

In today’s competitive B2B marketing landscape, identifying high-performing campaigns before they launch is critical. Traditional campaign reporting often relies on hindsight, leaving marketers reactive instead of proactive.

RevSure’s Campaign Performance Prediction Engine is an AI-powered solution that analyzes historical and real-time campaign data to predict future outcomes. It helps B2B marketing and revenue teams identify high-impact opportunities, allocate budgets efficiently, and adapt campaigns quickly for maximum performance.

By unifying data from multiple channels and providing actionable recommendations, it enables teams to:

  • Forecast campaign success before launch.

  • Optimize active campaigns to ensure they deliver results.

  • Improve marketing ROI through smarter resource allocation.

  • The model leverages historical performance and attribution data from the existing campaigns.

  • It analyzes key attributes such as,

    • Campaign Spend

    • Campaign Budget

    • Campaign Type

    • Campaign Name

    • Length of Campaign

    • Campaign Sources

    • Campaign Teams

    • Campaign Role

    • Campaign Start Date

    • Campaign End Date

    • Campaign Source Created At

    • Campaign Source Updated At

    • Generated Pipeline Value

    • Generated Pipeline Volume

How It Works

  • The framework uses advanced machine learning models to forecast pipeline/booking metrics like potential volume and potential value for upcoming campaigns.

  • Data Collection and Preparation

    • Features Utilization:

      • Campaign features such as campaign length are derived.

      • Date-based features like month, quarter, and year are extracted from date fields.

      • Audience and role-related features are incorporated, such as team or paid ads data.

      • Pipeline Generation Metrics: Volume and value metrics on the day of campaign start.

      • Historical averages of Open Pipeline Value and  Open Pipeline Volume.

      • Active campaigns at a given time.

      • Role-based contributions

      • Historical averages for similar campaigns.

  • Machine Learning Models

    • Two-Step Modeling Process:

      • The framework employs a two-step approach:

      • Classification Model:

        1. Predicts the likelihood of a campaign generating meaningful volume.

        2. Outputs a binary classification (high-potential or low-potential).

      • Regression Model:

        1. Predicts the potential pipeline/booking value or volume based on the output of the classification model.

    • Model Types:

      • Classification:

        1. Algorithm: Uses advanced algorithms like XGBoost Classifier.

        2. Target: Binary target indicating whether a campaign is likely to exceed a certain pipeline threshold.

      • Regression:

        1. Algorithm: Regression models such as XGBoost Regressor.

        2. Target: Continuous target representing pipeline value or volume.

Retraining Frequency: Once every quarter, or earlier if there are changes in data configurations.
Scoring Frequency: Daily

Key Features

Demand Generation Effectiveness

  • Predicts campaign pipeline and revenue performance

  • Recommends the best campaigns to double down on

Funnel Conversion Attribution

  • Captures contributions and role of campaigns  across Marketing, SDR/BDR, and AE motion in driving conversions at different stages of the funnel

  • Recommends the top journeys to orchestrate to drive conversions.

Deep Funnel Attribution

  • Quantifies campaigns' contributions to pipeline and revenue.

Benefits of Using RevSure’s Prediction Engine

Proactive Campaign Management

  • Predict outcomes before launch to ensure smarter planning.

  • Stay ahead of issues with real-time adjustments.

Increased Marketing ROI

  • Focus spending on campaigns that deliver the highest returns.

  • Avoid wasting budget on underperforming channels.

Faster Decision-Making

  • Act on real-time insights instead of waiting for post-campaign reporting.

Improved Resource Allocation

  • Optimize team efforts and budgets to focus on strategies that drive measurable results.

Cross-Channel Visibility

  • Get a unified view of all marketing campaigns, enabling better collaboration between marketing, sales, and RevOps.

Data-Driven Confidence

  • Eliminate guesswork by relying on AI-backed insights for every decision.

RevSure’s Campaign Performance Prediction Engine provides B2B marketers with the tools to forecast, monitor, and optimize campaign performance with AI precision. Delivering predictive insights and actionable recommendations ensures every campaign is positioned for success.

With its real-time monitoring, multi-channel tracking, and resource optimization capabilities, the engine enables teams to stay proactive, maximize ROI, and drive measurable results.

Marketing Mix Modeling Methodology

RevSure’s marketing mix modeling (MMX)  solution helps CMOs and their marketing teams make better planning and spend allocation decisions across different channels and investments such as  Dark Social (Podcasts, etc.), Brand, Paid Ads  (Google, LinkedIn, Meta, Bing), Events, Sales investments, Organic Search, Content and OOH.

MMX uses a multi-variate AI-based regression engine that quantifies the contribution and ROI of each channel towards pipeline and bookings for every quarter in the past and on a rolling basis. It can account for the impact of macro-economic factors and competitive events and actions.

It further provides the ability for marketing teams to integrate and quantify the impact of dark social and brand awareness metrics on overall long-term and short-term pipeline and bookings performance.

MMX insights enable customers to determine the most effective marketing budget allocation across channels for the upcoming quarters.

See data science overview deck here.

Key Steps and Principles:

The below steps are utilized to generate MMx contributions for Generated Pipeline Volume, Pipeline Value, Booked Volume, and Booked Value, as well as intermediate funnel stages (as desired)

Note: In the procedure below, predicted sales are used to represent both generated value and generated volume.

  • Data Preparation:

    • Historical marketing, sales, and macroeconomic data are gathered to provide a comprehensive view.

    • Intuitive and relevant features across channels like Facebook, Google, LinkedIn, Email Marketing, and Sales activities are engineered.

    • RevSure’s methodology also integrated touch based and time varying life analysis features to anchor the marketing mix measurement with the B2B buyer journeys.

    • Features are grouped into meaningful marketing categories to simplify interpretation and enhance decision-making. E.g.:

      • LinkedIn: LinkedIn spend, impressions, engagements, campaign count, audience size, audience profile

      • Google: Google spend, impressions, engagements, campaign count, audience size, audience profile

      • Facebook: Facebook spend, impressions, engagements, campaign count, audience size, audience profile

      • Events: Event spend, event count, audience size/profile

      • Organic Social: Organic posts, impressions, engagement (clicks, shares, reactions, comments), etc

  • Establishing Baseline Sales:

    • A strong baseline using natural business growth trends and seasonality is built, ensuring fair attribution of marketing effects.

    • Marketing contribution is measured only above and beyond organic growth.

  • Model Training and Feature Selection:

    • Mutual Information is applied to select the most influential features for each quarter and target variable.

    • Bayesian Ridge Regression, known for delivering stable and reliable results even in high-dimensional scenarios, is used.

    • The model is trained with K-Fold Cross-Validation to enhance robustness.

Feature Selection: Mutual Information, PCA

Retraining Frequency: Once every quarter during Projection Engine training, or earlier if data configurations change.
Scoring Frequency: Daily/Weekly (depending on the deployment)

How Marketing Mix Contribution is Calculated

  • For each marketing feature:

    • We create a counterfactual by neutralizing the feature (setting it to its baseline value) while keeping other features constant.

    • We predict the sales outcome with this modified feature.

    • The contribution is calculated as the difference between actual sales and counterfactual predictions.

  • Most Probable Sales Calculation:

    • We use a PCA-based conditional distribution sampling approach.

    • It simulates a realistic sales prediction if the specific feature’s influence is "removed."

  • Final Scaling:

    • All feature contributions are scaled proportionally to sum exactly to the total predicted sales for each record.

    • This maintains trust, alignment, and completeness in the modeling output.

How Response Curves are Built

  • For each marketing input:

    • The feature's value is gradually varied across a realistic and business-relevant range.

    • At each point, the corresponding Response Metrics are computed.

    • A smooth, best-fit curve (e.g., Hill Curve, Sigmoid, Power Curve) is generated.

  • Elasticity and ROI Insights:

    • The final curve shows how sensitive sales are to marketing investments.

    • It helps identify points of diminishing returns.

    • It enables ROI calculations (Return on Investment) and optimizes spend allocation.

Difference Between Marketing Contribution and Response Curves

Marketing Contribution:

This answers the question:

“How much did each marketing activity actually contribute to sales during a specific period, given what we observed?”

  • For each marketing feature (e.g., Facebook spend), we estimate its contribution by simulating a range of realistic scenarios in which that specific feature is held fixed. The remaining features are re-sampled based on the learned relationships from the data (via PCA and multivariate Gaussian distribution).

  • This allows us to generate a counterfactual prediction, i.e., what the model would have predicted if that feature had remained at its baseline influence, while the rest of the feature context stayed consistent.

  • The difference between the original prediction (with all features active) and this counterfactual is taken as the contribution of that feature.

Note: For past quarters, we have the actual value, and for the latest day of the running quarter, we predict value using the best-trained model.

  • We repeat this process one feature at a time to isolate the individual lift of each marketing variable.

  • Finally, contributions are scaled to ensure they sum exactly to the predicted sales value for each record, preserving both accuracy and interpretability.

Response Curves:

This answers the question:

“What would happen to sales incrementally if we changed the level of investment in a particular marketing activity, even outside the historical range?”

  • It is forward-looking and exploratory in nature.

  • Instead of simulating a "removal," we intentionally vary a specific marketing input (e.g., LinkedIn spend) across a realistic range, from very low to very high values.

  • At each level of this input:

    • We use the trained model to predict the corresponding sales outcome.

    • To ensure realistic results, other features are either held constant or allowed to adjust based on their historical correlation.

  • These prediction points are collected and used to fit a smooth best-fit curve (e.g., power, logarithmic, or sigmoid) that reflects the response behavior of that marketing input.

  • This curve shows whether doubling spend would double sales (linear), provide diminishing returns (concave), or show no effect.

  • It is used to understand elasticity and help with budget planning and optimization.

How to Interpret MMX Contribution vs. Response Curve

1. Marketing Mix Contribution

“What was the actual impact of LinkedIn spending on today’s pipeline?”

  • This is a backward-looking measurement.

  • It quantifies how much LinkedIn (or any channel) contributed based on what was already spent and what actually happened.

  • It's calculated in the context of other features and actual performance.

  • It is relative to the current day, using observed inputs and model attribution.

Use this when you want to explain what drove performance historically.

2. Response Curve

“If we hypothetically spent $24K on LinkedIn by the end of the quarter, what pipeline value would we expect?”

  • This is a forward-looking tool used for planning and simulation.

  • It evaluates how much pipeline could be generated if we change spend, often well beyond what’s been observed so far.

  • It assumes all other features stay constant.

Use this for forecasting, budgeting, and understanding ROI trade-offs.

MMX Contribution = What happened given the actual inputs
Response Curve = What could happen if we change our spending going forward

Statistical Incrementality & Lift Analysis

Incrementality Testing is a statistical approach that helps marketers statistically measure the lift or additional impact of a marketing campaign or channel, or treatment.

Traditional rules based attribution models often give credit to all touchpoints (multi-touch) or just one (first/last touch), but they are not backed by statistical measurement. Incrementality & Lift Analysis closes that gap by providing confidence-backed answers to these key questions:

  • Did this campaign influence an outcome with statistical measurement and validation?

  • What’s the measurable lift from this paid channel?

  • Should we scale this initiative, or sunset it?

RevSure supports the following Incrementality & Lift Analysis methods and widgets.

  • Conversion Lift Analysis

  • Pre-Post Analysis

  • Multi-variate Causal Impact Analysis

  • Campaign Lift Analysis - Across Campaigns

  • Campaign Lift Analysis - Individual Campaigns

Conversion Lift Analysis Widget

This widget helps measure the statistical validity of conversion differences between two groups:

Control (where a particular marketing / GTM activity has not been performed)
Test (for which a particular marketing/GTM activity has been performed)

In RevSure, you can identify Control vs. Test based on Boolean (true/false) or Binary Categorical (Y/N) dimensions, such as (Y/N). Examples of such dimensions include Lead Google Touched and Lead LinkedIn Touched, among others.

The statistical test is done using a Chi-squared-T test for proportions.

Minimum Sample Size Check
Before running the test, we calculate the minimum required sample size per group using a Chi-Square.

  • Inputs: expected control/test conversion rates, significance level, and desired statistical power.

  • This ensures that the groups are large enough for the Chi-Square approximation to be valid and that the test has sufficient statistical power.

  • If the observed sample size is below this threshold, results may be flagged as unreliable.

The statistical significance is indicated by the p-value of the test:

If the p-value < 0.05, then the Lift is considered statistically significant. This means, the inference can be made that the marketing/GTM activity made a statistically significant impact on the conversions.

Else if the p-value > 0.05 the Lift is considered statistically insignificant. This means that no inference can be made on whether the marketing/GTM activity made an impact on the conversion. Any lift or difference observed could just be a random observation.

Pre & Post Analysis Widget

In RevSure, the ‘Pre & Post Analysis’ widget in the Marketing Performance module is designed to help with testing the impact of a marketing/GTM Treatment on a Response Variable.

The Treatment could be running a new campaign, an Event, a competitor action, increasing digital spend, etc.

The Response could be an outcome such as pipeline generation, lead generation, or booking value generation. etc.

This analysis is designed to statistically compare the difference in the Response variable before the treatment (pre) and the Response variable (post) the treatment.

The statistical method used in this test is the 2-sample Welch t-test.

Minimum Sample Size Method used: Two-sample one-tailed t-test–based sample size calculation with finite population correction (FPC) method is used. This method also includes calculations of the cohen’s distance, which helps to identify how far two samples lie based on their mean and variance.

  • One-tailed: because it assumes directionality (e.g., post ratio > pre ratio).

  • Two-sample: because it compares variance from pre-period vs post-period ratios.

Statistical Significance of the Pre-Post Test:

The statistical significance is indicated by the p-value of the test.

If the p-value < 0.05, then the Difference and Outcome are considered statistically significant. This means the inference can be made that the Treatment made a statistically significant impact on the Response.

Else if the p-value > 0.05, the measured Difference and Outcome are considered statistically insignificant. This means that no inference can be made on whether the Treatment made an impact on the Response. Any difference observed/measured could just be a random observation.

Difference-in-Differences (DiD) Analysis

The Difference-in-Differences (DiD) Analysis in RevSure is designed to measure the true incremental impact of a marketing or GTM initiative by comparing outcomes across Test and Control groups over time. It builds on the Pre & Post Analysis by introducing a Control group to isolate the effect of the treatment from broader market or seasonal trends.

What is Difference-in-Differences?

Difference-in-Differences compares how a key business outcome changes over time for:

  • A Test group that received the treatment

  • A Control group that did not receive the treatment

By comparing these two changes, DiD removes external factors that affect both groups and isolates the impact directly attributable to the treatment.

Both Test and Control groups are typically influenced by common external factors such as seasonality, macro trends, or market shifts.

DiD answers: Did the Test group improve more than the Control group after the treatment?
If yes, the additional improvement is considered the true incremental impact of the initiative.

Key Concepts

  1. Treatment: The intervention being evaluated.
    Examples: campaign launch, increased paid spend, event execution.

  2. Response Variable: The outcome influenced by the treatment.
    Examples: pipeline generation, lead generation, booking value.

  3. Test Group: Entities that received the treatment.

  4. Control Group: Entities that did not receive the treatment and act as a baseline.

How RevSure Measures Impact
RevSure evaluates the change in the Response Variable for both groups:

  • Change in Test group (Post vs Pre)

  • Change in Control group (Post vs Pre)

The Difference-in-Differences effect is computed as: (Change in Test group) − (Change in Control group)
This represents the incremental lift driven by the treatment, after removing underlying trends.

DiD Effect (Lift)
The estimated incremental impact of the treatment.

  • Positive value = Treatment drove improvement beyond baseline trends

  • Negative value = Treatment underperformed relative to baseline

  • Near zero = No measurable incremental impact

Statistical Significance
RevSure evaluates whether the observed lift is statistically reliable:

  • A two-sided t-test is applied to the DiD coefficient

  • Uses classical OLS standard errors

  • Determines whether the observed impact is statistically significant

Multi-variate Causal Impact Analysis

This version is an advanced version of the Pre-Post Analysis, which includes a more rigorous causal impact analysis along with the ability to analyze the impact of marketing and GTM interventions that might involve multiple levers/tactics/campaigns.

The methodology is inspired by Google’s Causal Impact approach and data science methodology:
https://google.github.io/CausalImpact/CausalImpact.html

It takes two datasets:
Pre-period: before the treatment/campaign/change.
Post-period: after the treatment/campaign/change.

  • We pass one response/KPI (target), a list of covariates (variables not affected by the treatment/campaign), and optional treatment fields (kept for reporting).

  • We fit a time‑series model that learns how the KPI normally moves with several covariate variables. Training on the pre-period prevents the treatment’s or campaign’s effect from leaking into the model.

  • In the after window, we plug in the variables, and the model forecasts what the KPI (Counterfactual forecast) would have been without the initiative. The forecast includes a confidence band (uncertainty).

Post that the impact is computed:

  • Compute impact, Impact per day/week = Actual - Counterfactual; we also show cumulative lift and % lift.

  • We report a Bayesian equivalent of the statistical significance p-value of whether the lift > 0 to convey confidence.

  • If the p-value < 0.05, then the Impact/Lift is considered significant..

  • Else if the p-value > 0.05, the Lift/Impact is considered statistically insignificant. This means that no inference can be made on whether the marketing/GTM activity made an impact on the lift.  Any lift or difference observed could just be a random observation.

Campaign Lift Analysis - Across Campaigns

The Campaign Lift Analysis widget helps you measure the impact of specific campaigns, campaign types, or channels on conversion rates across your funnel. It uses statistical significance to determine whether a campaign meaningfully contributes to conversion from one stage to another.

This analysis compares the conversion lift from a selected campaign against others using a Chi-Squared test.

Understand the Output:
Conversion %:
Percentage of leads/accounts with a touch of that campaign/channel/type in the journey from the start stage to the end stage that converted to the end stage.

Lift %: How much higher (or lower) the conversion % is compared to the base.

Minimum Sample Size Method used: Two-proportion z-test

Statistical Significance:

If a campaign/channel shows a Significant Lift, it’s backed by a p-value < 0.05, indicating high confidence.

Example: Marketing Form-Fill shows +1903.57% Lift with a strong p-value – its first touch significantly boosts conversion (72.73%) from Lead to MQL against all other campaigns (3.63%)

Calculation Logic of Campaign Lift Analysis Across Campaigns

  1. Calculating the conversion % of a campaign

    1. Calculate the Number of leads with a campaign touch that converted - X

    2. Calculate the Number of leads with a campaign touch that did not convert - Y

    3. Conversion % of a campaign = X/(X+Y)

  2. Calculating the conversion % of all campaigns

    1. Calculate the Number of leads with any of the campaigns touched that converted - X

    2. Calculate the Number of leads with any of the campaigns touched that did not convert - Y

    3. Conversion % = X / (X+Y)

  3. Calculating the conversion % of selected campaigns

    1. Calculate the Number of leads with any of the selected campaigns touched that converted - X

    2. Calculate the Number of leads with any of the selected campaigns touched that did not convert - Y

    3. Conversion % = X/(X+Y)

  4. Lift is calculated,

    1. When comparing against a base value of the average of all campaigns, (Conversion of a campaign- Conversion of all other campaigns)/(Conversion of all other campaigns)

    2. When comparing against a base value of the average of selected campaigns, (Conversion of a campaign- Conversion of selected other campaigns)/(Conversion of selected other campaigns)

    3. When comparing against a base value of a campaign, (Conversion of a campaign- Conversion of base campaign)/(Conversion of base campaign)

Campaign Lift Analysis - Individual Campaigns

The Campaign Lift Analysis - Individual Campaigns widget uses the same methodology as the Campaign Lift Analysis - Across Campaigns widget. The difference lies in how the lift is calculated

Campaign Lift Analysis - Individual Campaigns compares lift of the campaign w/o reference to any other campaign/touch. It analyzes whether having the touch of the campaign is more beneficial than not having it. For example, in the case of first touch, it analyzes whether an MQL with the first touch of Marketing Form-Fill has a higher chance of conversion than not having the first touch of Marketing Form-Fill

Minimum Sample Size Method used: Two-proportion z-test

Calculation Logics of Campaign Lift Analysis - Individual Campaigns

  1. Calculating the conversion % of a campaign touch

    1. Calculate the Number of leads with a campaign touch that converted - X

    2. Calculate the Number of leads with a campaign touch that did not convert - Y

    3. Calculate the Number of leads without a campaign touch that converted - X'

    4. Calculate the Number of leads without a campaign touch that did not convert - Y'

    5. Conversion % of True = X/(X+Y)

    6. Conversion % of False = X'/(X'+Y')

  2. Calculating the lift

    1. If Control Value is False, (Conversion of True - Conversion of False)/Conversion of False

    2. If Control Value is True, (Conversion of False - Conversion of True)/Conversion of True

Closing Note

While this document provides an overview of each model, the true value of RevSure’s AI framework lies in how seamlessly it integrates into day-to-day GTM workflows. These models aren’t just analytical engines, they are decision enablers built for RevOps, Marketing, and Sales teams to act with clarity and precision. From campaign planning to pipeline reviews, RevSure helps teams move from lagging indicators to leading actions.

As buyer behavior, channels, and data patterns evolve, so will these models. Continuous learning and adaptation are core to our approach, ensuring that your GTM engine becomes not just data-driven, but self-optimizing over time.