Most organizations implement an operating model based on frameworks like Objective and Key Results (OKR). Teams establish measurable targets, execute specific initiatives designed to achieve them, and repeat the cycle every quarter.
This exercise is great for aligning teams and quantifying outcomes, but it often misses a critical step: measuring the efficacy of initiatives themselves and using that data as a feedback loop for subsequent decision making.
A standard operating model might look something like this: Leadership sets annual thematic goals. Each quarter, departments align with these themes to set target metrics and define the key initiatives that will influence each metric. For example, if an annual theme is growth, a sales team may set out to close a $10M bookings target, establishing 3 key strategies they believe will help achieve this goal. At the end of the quarter, the organization reviews the outcomes to understand how they performed. If the sales team closed $8M in bookings, they achieved 80% of their goal.
Visually, this looks something like this:
This is precisely where many organizations stop, choosing a new set of initiatives without completing the feedback loop. Disciplined teams may discuss the impact of each initiative at a high level, but rarely in a quantifiable way. Failing to look closely at the underlying drivers of each outcome presents two main issues:
All initiatives, no matter how they performed, are measured against a single target metric. If our sales team achieved 80% of its target, which initiatives actually drove positive results and which were detractors? Of this distribution, what was the ROI for each?
Without a quantifiable understanding of each initiative, the quarterly cycle restarts and the feedback loop is lost. Yes, an outcome was achieved, but we failed to internalize the most valuable data: the key drivers of each initiative that lead to the outcome.
Without comprehensive measurement, an operating model is limited to serving as a lagging indicator. It does an excellent job at telling us how we performed and a terrible job at telling us how to perform better in the future. At a human level, this sort of deep retrospection doesn’t always come naturally.
The human lens: Why do we have a blind spot for learning how to do better next time?
It’s a double whammy, actually. Most of us (unless we’ve had MUCH therapy) avoid the truth that we are imperfect. We all have an inherent bias to view ourselves as “great” because anything less feels like a threat to our status and belonging. Our brains are wired to solve for status and belonging, and one of our default strategies is self-deception.
If that weren’t enough, the vast majority of us were raised in fixed mindset systems – systems that rewarded us for having the right answer (of which there was only one). High performers who got straight As were trained to solve for the outcome, not the learning.
The challenge is that the world has changed much faster than our systems and our evolutionary biology. In a world of endless information, constant change, and vast uncertainty, we need to completely reorient our mindset. We must embrace curiosity to accept that there may be many “right” answers. It’s our job to discern which one is best for our context and take action with skeptical curiosity. We need to be focused on GETTING it right, even if that takes multiple iterations, rather than BEING right. It requires that we view ourselves as “continuous learners” as opposed to “subject-matter experts.” And that can feel like a hit to both our power and our status.
What can technology teach us about goal-oriented feedback systems?
Deep Learning is a subset of Machine Learning and AI that uses artificial neural networks to learn hierarchical representations of data — and it does this in a fascinating way.
These models use a concept called backpropagation. During training, the model makes a prediction, and this prediction is compared to a ground truth – an answer key of sorts. Given the accuracy of each prediction, the model traverses back through its layers, using backpropagation to isolate the error contribution of each layer. An optimizer then adjusts the model’s weights and biases in an effort to be “less wrong” in the future. This process is repeated millions of times until its loss score (the mathematical measure of how far predictions deviate from actual target values) reaches an acceptable minimum.
In other words, the model has a goal, it makes a prediction, and then it doesn’t simply report its performance. Instead, it methodically examines how each individual component contributed to the positive or negative outcome and adjusts its “behavior” accordingly. Each of these micro-experiments compound quickly, pushing the model towards incrementally better outcomes.
This isn’t a new concept for humans; we perform the same type of calculations on a daily basis to adjust our own behavior patterns, though often unconsciously. In Psycho-Cybernetics, published over 60 years ago, the author Dr. Maltz describes this same self-learning feedback mechanism for humans:
“All skill learning is accomplished by trial and error, by making a trial, missing the mark, consciously remembering the degree of error, and making corrections on the next trial–until finally a hit, or successful attempt, is accomplished. The successful reaction pattern is then remembered, or recalled, and imitated on future trials.”
What does it mean to be “less wrong” from a human perspective?
What if we borrowed the goal of being incrementally “less wrong” from deep learning? If we started from the premise that there is no way to get it “right” the first time? With that framing, hitting 80% of our goal is just a data point – the metric we need to trigger the question, ‘What would we have had to do differently to get to 100%?’ (focus on learning) vs. ‘I fell short of my goal by 80% so I’ve failed - I’d better get my defenses ready as to why 80% is still good enough’ (focus on status).
If we can admit that setting an OKR is just an educated guess, then our work is just to become “less wrong” each quarter when we review key results. At its core, that’s how the OKR framework was created: ”stretch” goals were explicit guesses.
The rub comes when we tie performance reviews and compensation to results. If my bonus or raise is based on achieving a number, it’s almost impossible for my status and power-focused ego to leave any room for learning; I’m just hyper-focused on the outcome that affects me directly. While decoupling performance evaluation and goal-setting is beyond the scope of this article, I’d offer these two questions for discussion in your own organization or context: (1) Which mindsets and behaviors is our current system incentivizing? (2) How might we align our incentive system to support the shift to a learning mindset?
The Learning-Oriented Operating Model
Of course, an organization can’t be expected to trace every incremental action that leads to an outcome. Nor do we want to introduce a mountain of unnecessary administrative work. However, it can and should measure the efficacy of each initiative in detail.
Objective
Our operating model shouldn’t simply measure outcomes. It should have a deep understanding of what influenced each outcome, such that we can adjust behavior in a systematic effort to improve future outcomes.
Prerequisites
Moving from static outcomes to an operating model with a quantifiable feedback loop depends on three key prerequisites:
Adjusting the operating model to adopt a hierarchical structure in which each target metric has 3-5 initiatives and each initiative has 3-5 key drivers.
Ensuring the metrics attached to each layer are quantifiable and can be normalized to a relative percentage score.
Implementing systemized “lab logs” that ensure every important action or decision can be recorded and attributed back to a driver and its initiative. To reduce administrative burden, it’s recommended to rely on a combination of existing action logs (like actions taken in a CRM) and AI tools where teams can “brain dump” their decisions and findings along the way.
Since every organization is nuanced, it’s encouraged to adopt these concepts in a way that best fits your team’s structure and desired operating model.
Definitions
Our new operating model should have three key components:
Target Metric: A high-level goal the organization or department is trying to achieve. For example, achieving a bookings target of $10M, the deployment of a product roadmap, or improving customer retention by 10%.
Initiatives: The high-level, tactical activities that will move the target metric in the right direction. For example, in order to achieve a bookings target of $10M, we may need to increase our average deal size by 35% while improving lead velocity by 10%.
Drivers: The specific experiments we will conduct, levers we will pull, or variables we will try to influence. If our initiative is to increase the average deal size by 35%, collectively, what measurable experiments do we hypothosize will drive our desired outcomes?
While both initiatives and drivers receive a quantifiable metric, its important to distinguish that initiatives are measurable results while drivers are simply the result of our experimentation. For example, if a driver hits 133% of its target but the parent initiative fails, our conclusion should not be that our experiment succeeded but rather, that our hypothesis was wrong. The driver didn’t influence the initiative’s metric.
Let’s return to our sales team who had a $10M target and three planned initiatives. Their new quarterly hypothesis might look something like the example below. Note, we are not sales leaders and the examples below are purely illustrative.
Illustrative Quarterly Plan
Organization Theme
Growth
Target Metric
$10M Booked ARR
Initiatives
Initiative 1: Up-Market Migration → Moving from mid-market deals to larger enterprise accounts to increase average deal size by 35%.
Driver 1A → Time to Enterprise Readiness: Implement a multi-stage certification program (shadowing, pitch certifications, and deal-desk sign-offs) to measure the average days required for an AE to independently pitch, manage, and close $100k+ ARR deals without executive shadowing, moving from 45 days to 20 days.
Driver 1B → Multithreaded Stakeholder Rate: Require reps to map and log executive contacts during discovery, tracking the percentage of open qualified pipeline deals with at least 4 verified decision-makers/economic buyers actively logged in the CRM.
Driver 1C → Custom PoC Conversion Rate: “PoC Evaluation Criteria” sign-off between Sales Engineering and the prospect before building, measuring the percentage of completed technical proofs-of-concept that successfully convert to formal commercial proposals.
Initiative 2: Improve Inbound Lead Velocity → Maximizing revenue output from incoming marketing-generated pipeline by 10% without increasing headcount.
Driver 2A → Speed-to-First-Touch: Deploy automated round-robin lead routing and Slack/SMS rep notifications to measure the median time (in minutes) from an inbound form submission to a rep’s initial manual outreach or completed meeting booking. Target 20 minutes.
Driver 2B → MQL-to-SQL Acceptance Rate: Enforce a 24-hour SLA and mandatory disposition logging in the CRM for reps to either accept or formally reject marketing-qualified leads with a structured drop-down reason, tracking the percentage of MQLs accepted into active sales sequences.
Driver 2C → Discovery-to-Pipeline Conversion: Implement a standardized BANT/MEDDPICC qualification checklist in discovery calls to track the percentage of completed first meetings that successfully convert into formal pipeline opportunities with an agreed-upon target close date.
Initiative 3: Expanded Cross-Selling → Expanding existing customer revenue by an average of 5% through upsells, add-on modules, or seats.
Driver 3A → Early Account Health Score: Execute a 90-day structured customer onboarding playbook (kickoff call, core integration, and initial team training) to measure the percentage of newly signed accounts reaching “Green” health status across usage, support ticket, and sentiment metrics before day 90.
Driver 3B → Executive Touchpoint Frequency: Establish a recurring quarterly executive sponsorship program that pairs internal leadership with client counterparts, tracking the average number of completed executive-level strategic alignment syncs per account each quarter.
Driver 3C → Product Adoption Depth: Automate targeted product-nudging campaigns and Customer Success check-ins triggered by inactivity, tracking the ratio of licensed seats active weekly (WAU/MAU) and the adoption rate of core high-value features.
The key drivers that previously sat below the surface (with low visibility but high impact) have now become the foundational anchor of our operating model. Visually, this transformation looks something like this:
As the quarter kicks off, each team executes against their initiatives, performing a series of rapid, iterative experimentation to influence each driver. By the end of the quarter, we’ll know how each metric, initiative, and driver performed – and we’ll have a detailed history of why by way of our lab logs.
The Art of Attribution
Before we can ever say with confidence which initiatives and drivers impacted our target metrics, we must first define the degree of the relationship or independence between each. This requires us to consider whether our target metric is a composite of its underlying initiatives and drivers or if each should be measured as an independent diagnostic.
The ambiguity of art and science that exists in attribution is a reason many organizations abandon (or fail to even start) this practice. It’s not always a simple, neat construct. But, working towards grounded attribution can unlock an entirely new world of understanding and operational rigor.
Composite Attribution
In composite attribution, drivers are defined such that their metrics “roll up” (via sum, weighted average, or some explicit formula) to produce their parent initiative and target metric. The target metric still receives a target or goal, but the outcome is a derivative of its underlying components rather than an external measurement (like dollars booked in sales).
This approach is particularly useful in two scenarios:
When initiatives and drivers are the only mechanics capable of moving the target metric. For example, in a Product & Engineering organization, if our target metric represents the delivery of a quarterly roadmap, the initiatives themselves directly roll up to achieving this target metric.
When our objective is to incrementally improve a target metric through experimentation. For example, if a customer enablement team is working to improve a days-to-launch metric, seeing the target metric influenced directly from ground-floor drivers can have an immediate, positive impact.
In composite attribution, organizations can also get creative with different types of weighted calculations. For example, if 2 of the product roadmap deliverables have an outsized impact on a core business metric, these can be weighted when producing the final target metric, ensuring teams were properly allocated towards the highest ROI activities.
Diagnostic Attribution
Diagnostic attribution, on the other hand, is ideal when our target metric is a grounded, core business objective with little ambiguity. For example, momentarily remove all initiatives and drivers: our target metric for sales bookings was $10M and we achieved $8M; thus, an 80% outcome. Now, our job is to evaluate the diagnostic signals of each initiative and driver to understand how each may have impacted our target metric.
In diagnostic attribution, each initiative and driver is measured on its own, against its own definition of success. Since every level is independently measured, driver metrics don’t sum to their parent initiative’s metric, and initiative percentages don’t sum to the target metric. They’re not supposed to. Drivers and initiatives in this case are diagnostic. They explain the outcome, they don’t mathematically produce it.
Since driver and initiative metrics serve as hypotheses rather than mathematical components, leaders must treat them as diagnostic signals rather than absolute proof.
These signals must be paired with qualitative evidence from lab logs, leadership experience, and other factors. During post-quarter reviews, leaders shouldn’t just look at whether a driver hit its percentage goal; they should cross-reference performance with the team’s lab logs to ask: Can we trace specific outcomes directly back to the actions taken here, or did an external tailwind do the heavy lifting?
Using diagnostic drivers in this way gives leaders directional accountability without falling into the trap of false mathematical precision. It allows teams to refine their operational hypotheses quarter after quarter, getting progressively “more right” about which levers actually move the needle.
Initial Results
Let’s look briefly at an example of what we are actually trying to extract from this exercise before diving too deeply into the calculations. Despite our attribution method, each target metric, initiative, and driver has a quantifiable outcome — whether a composite or independently measured variable. For our sales example, we’ll use diagnostic attribution.
Target Metric: 80% ($8M of $10M Booked)
Initiative 1 → 115%
Driver 1A → 133%
Driver 1B → 121%
Driver 1C → 90%
Initiative 2 → 60%
Driver 2A → 33%
Driver 2B → 70%
Driver 2C → 95%
Initiative 3 → 85%
Driver 3A → 50%
Driver 3B → 90%
Driver 3C → 100%
The metric for our first initiative was to increase deal size by 35% and we’ve hit just over 39% – surpassing our target. Initiative 2, on the other hand, performed poorly and initial diagnostic signals point Driver 2A.
We can now begin to evaluate our bookings miss at a deep level, pinpointing what worked and what didn’t. Just as importantly, we can now put these outcomes in context. For example, we might note:
Driver 1A performed really well, but it actually wasn’t until the latter half of the quarter that we saw the needle move. What experimentation catalyzed this shift?
Initiative 2 was our poorest performer, dragged down primarily by Driver 2A. We’ve seen some performance issues on this team and their lab logs suggest they need retraining in a particular area of the business.
Initiative 3 performed OK, but Driver 3A missed the mark. The lab logs are comprehensive, but we can see that the team was hung up on a particular strategy that clearly wasn’t performing well and they failed to evolve that strategy over the quarter.
This level of attribution, paired with AI tooling, allows leaders to develop a deep understanding of their team’s performance and provide clear, data-driven guidance for the next quarter.
Running the Numbers
Importantly, whether we use composite or diagnostic attribution, the calculations are the same. In a composite model, these calculations compute exact structural contribution. In a diagnostic model, this math serves as a normalized sensitivity score — a way to rank-order operational drag so leaders know where to best allocate diagnostic time.
There are three steps to go from statically measured outcomes to a weighted attribution of each metric, initiative, and driver. In this example calculation, we’ll speak in terms of drag (detracting from the target metric) and lift (contributing to the target metric).
First, we want to measure the positive (lift) or negative (drag) movement of each individual driver: Achieved % - 100%. For example:
Driver 1A: 133% - 100% = +33 → Lift
Driver 1B: 121% - 100% = +21 → Lift
Driver 1C: 90% - 100% = -10 → Drag
Do this for each driver across all initiatives. Then, calculate the gross lift and gross drag by summing the absolute values for lift and drag separately. This gross calculation can be done per initiative or across the entire department.
Gross Lift: 33 + 21 = 54
Gross Drag: |-10| = 10
These gross values provide the denominator needed to calculate relative contribution at the initiative level:
Driver 1A: 33/54 → Contributed 61% of the lift for Initiative 1.
Driver 1B: 21/54 → Contributed 39% of the lift for Initiative 1.
Driver 1C: 10/10 → Contributed 100% of the drag within Initiative 1.
If we apply this same concept across all initiatives in the department where Total Lift = 54 and Total Drag 172, we can evaluate cross-departmental impact. We can see that Driver 1C appears to have contributed to 5.8% of the drag across all initiatives. Or, that Driver 3A may have been responsible for 29% of drag across all initiatives, even though its parent Initiative 3 performed 85% to target.
Without this perspective, we may have simply assumed that 85% was a reasonable outcome for a stretch goal without realizing the negative impact of one individual driver may have had. Instead, we can now double-click on this driver to either better understand its direct impact (in the case of composite attribution) or use these signals to perform diagnostic attribution.
Troubleshooting
In diagnostic attribution, variables do not necessarily have a linear influence. It’s important that leaders apply their understanding of the operational model to weigh influence accordingly.
Do not allow lab logs to become an operational burden. Lab logs are conceptual; they represent capturing the decision making processes surrounding drivers and initiatives — they are not a prescriptive, administrative requirement.
Conclusion
Many organizations fall victim to performative measurement where the outcome is clear, but the true underlying drivers remain hidden below the surface. When measurement stops at outcomes, we are actively discarding valuable data: what led to those outcomes and how to be “less wrong” with each subsequent experiment.
Leaders can instead embrace a Learning-Oriented Operating Model to establish target metrics, initiatives, and underlying drivers – each with clear, measurable outcomes. Depending on the team and what we’re measuring, these components can provide either composite or diagnostic level attribution. Throughout the quarter, teams experiment and iterate on ways to positively impact each driver. When the quarter is over, a quick series of calculations provides leaders with a comprehensive understanding of their organization’s performance. This granular data, lab logs, and an LLM then becomes a powerful tool for leaders looking to affect change across departments.
About the Authors
Erik Smith
Erik Smith is an entrepreneur, investor and builder. He serves as Chief Technology Officer at Rx Redefined, a venture-backed Series B healthcare technology company. His work centers on supporting leaders and teams across Product, Engineering, Data, AI, Business Intelligence, Security & Compliance, and IT/Infrastructure functions. He has a passion for helping to bridge the “curiosity gap” between technology, business, and operations. Erik lives in Northern California with his wife and two bunnies, Poppy and Bubbles.
If you’re interested in the intersection of business, technology and operations, follow Approachable X on Substack. Erik is working with Dr. David Cochran, Dean for AI and Emerging Technologies at Newman University, on their new book, Approachable AI for Business Leaders.
Jess Felts
Jess is Managing Partner at Team Level Partners, a boutique consulting firm focused on equipping leaders and teams with the human skills necessary to thrive in the era of AI. A connector at her core, Jess brings together her background as a strategy consultant, coach, and start-up leader to take an integrated approach to leadership and team effectiveness. She spent her early career in investment banking before attending the UC Berkeley Haas School of Business where she earned her MBA. She then went on to work at McKinsey, focusing on strategy and organizational consulting before shifting into the People space in the tech industry.
Often, the hardest problems in an organization aren’t technical, they’re relational: Trust that erodes under pressure. Conversations that avoid the real issue. Cross-functional friction that stalls execution. As AI reshapes how work gets done, the real differentiator is human skill: how people listen, build trust, adapt, and make high-stakes decisions together.
We’re expert facilitators, former operating leaders, and coaches with interdisciplinary backgrounds who work at the heart of how teams relate, decide, and move forward, especially when things get hard. If you’re interested in learning more about how to build the interpersonal muscles required for effective teaming and true organizational momentum, reach out to us at info@teamlevelpartners.com.








