AI tools like Figma Make have been shortening design cycles, handling the repetitive, mechanical work that slows teams down so they can focus on the work that matters most. Yet many teams aren't sure how much time they're actually saving, or where they're getting the biggest wins on efficiency. The Figma Data Science team set out to find the answers—but common research methods like A/B testing and causal inference couldn't account for all of AI's complexities.Overall, design work got 20% faster and 16% easier with Figma Make. PMs saw the biggest benefits: Tasks were 23% faster and 37% easier.So we designed a randomized controlled trial where we asked 100 participants to work through tasks with and without Figma Make. Here, we share more details about how we designed the study and the results we uncovered.Controlling for confoundersMeasuring something like time savings from AI usage is tricky because of confounders: other variables that also affect how long a task takes. Job tenure, career experience, and task complexity are all examples. Each can influence how quickly someone works, regardless of AI. Most everyday statistical methods can't account for confounders like these. So without controlling for them, we can't tell how much of a productivity change is really due to AI, versus other factors at play.Challenges with common methodsWhen it comes to confounders, two common data science methods—online A/B testing and causal inference on data logs—have their limitations.Imagine that we provided a test group of users with Figma Make and turned it off for another group. Then we would measure how long it takes group A to complete design tasks versus group B. But even though this theoretically randomizes for confounders, the main problem is that we can’t control the specific tasks each user is doing to make sure they’re equally complex.Causal inference is a set of statistical methods used to determine whether one variable actually causes a change in another, instead of just correlating with it.Alternatively, imagine that we wanted to look back through log data and see how much time on average users spend on design tasks when they use AI features versus when they do not. There are several causal inference methods we could use to try to determine how much AI usage drove the outcome, while controlling for confounders. But neither really works for our purposes.Propensity score matching (PSM) is a method for estimating causal effects from observational data. For each data point (in our case, each design task in the logs), you estimate the "propensity score": the probability it would have belonged in the treatment experience (in our case, AI usage), given its observed characteristics. But this relies on all confounders being included in the log data. That’s not the case for us; for instance, anonymized user IDs don’t tell us anything about someone’s subjective design experience.Instrumental variables (IV) is a method for estimating causal effects when a confounder can't be directly measured or controlled for. To use this method, we’d need to find a new variable in the data—an "instrument"—that affects whether someone uses AI. That instrument can’t affect task speed in any other way, and it has to be unrelated to the confounders. If both of those hold, we can use the instrument to isolate a slice of AI/non-AI usage that's essentially random with respect to confounders. The problem is that there’s no valid instrument for us—nothing in the log data nudges someone toward using AI features for reasons unrelated to their design experience or task.Choosing to run a randomized control trialRandomized control trials (RCTs) are the gold standard for valid causal impact measurement and the approach we landed on for this study. RCTs allow us to fully control for confounders at the start of data collection through combining random assignment, standardized task design, and trial moderation:Randomizing the participant population between treatment versus control ensures cofounders are divided approximately equally between the two groups.Having participants complete the same tasks eliminates the impact of different task attributes on our outcome variable of time to completion.Trial moderation from a trained research team mitigates further variability in the conducted work post group assignment, and protects against other potential statistical violations.We partnered with user researchers at and outside of Figma to create and conduct an RCT to find out how much time product designers and product managers save in everyday design work with Figma Make.We chose Make because of its popularity with Figma users and its wide range of applications. The more AI tools participants could use in the study, the more variables we’re introducing regarding what may contribute to time savings, so we wanted to limit the study to one tool.To determine how many participants to include in the study, we used effect sizes observed in peer studies in the industry, such as the Github Copilot RCTs. Then we used that approximated effect size to perform a power analysis, which led us to our final sample size of 100 participants, with 50 product designers and 50 product managers. Power analyses determine sample sizes required to detect statistically significant effects if they do occur within the study.We knew we would measure time savings for product designers—our power users of Make—but we were also curious about whether product managers (PMs) were saving time. We already knew from user research that Make has opened the door to design for PMs, and we wanted to better understand their experience.Designing the experimentThe core design of the experiment was to have treatment and control participants work through the same design tasks. The only difference between the groups was that treatment would have access to Make, while control would have no AI tool access.When designing tasks for participants to work through, our first consideration was familiarity. We needed our tasks to be widely understood by all participants, no matter what their actual design experience was. We landed on editing social media posts.We wanted to test a range of common design workflows in the study so we’d have multiple signals on Make’s impact across a variety of design tasks, helping us land on a more robust cumulative measure of time savings. We identified the themes of UI and appearance changes, creating new views in a base design, and adding interactability and responsiveness as key design workflows we wanted to test, creating three scenarios in total.Task 1: Convert to dark modeTask 2: Add a help option to a settings menuTask 3: Build a comment fly upFinally, and most challenging of all, our study tasks needed to land in the “Goldilocks” zone of not being too easy for product designers, but also not too hard for product managers. We ran internal pilots with Figma designers and product managers to test out our tasks with live participants. We learned what was and wasn’t working in our scenarios, and we redesigned the tasks themselves three times before landing on their final versions.Moderating the studyOur next challenge: We needed to train a team of moderators on how to run the study for 100 participants. Differences in moderation could become another confounder, so we developed a script that all moderators followed to guide participants through the study workflow in a standardized format.But our moderators all had different levels of Figma fluency—which meant the level of guidance they could offer participants who got stuck during the study would vary. We wrote out a troubleshooting guide that moderators could reference to support participants, and revised the guide as the study went on when new issues came up.The resultsIdentifying our participants, designing the experiment, and preparing for moderation unearthed complexities that required multiple iterations from the team to fix. But despite these challenges, all 100 participants successfully completed the study. We arrived at these key findings:Overall, participants saw a 20% reduction in cumulative time to task completion, 16% improvement in task ease, and 15% improvement in perceived Figma usability.Product managers reported consistently greater gains than product designers across all tested dimensions. The time gain back for PMs if they have Make access was 23% cumulatively across the study tasks, according to our interaction regression model. In the sum total of the experiment, a product manager using Make was almost as efficient as a product designer who didn’t use Make.The final key finding is regarding which type of design tasks enable time savings gains for product designers compared to PMs. PMs reported gains on our two easiest and quickest tasks in the study, but they did not experience statistically significant improvements on the most challenging and extensive task in the study. The reverse is true for product designers: They only report stat sig task time improvements on the most challenging task, and not on the two simpler scenarios. In other words, Make promotes faster zero-to-one and complex interaction design for designers, while for PMs, it serves as more of a baseline tool to enable quick design contributions.We used a combination of hypothesis testing and OLS regression modeling to analyze our results. All of the results described here are statistically significant unless noted otherwise.Because of our study design, we can confidently say that Make causes these improvements. This approach was robust against confounders and other forms of biases, and we plan to continue to use RCTs to answer top-of-mind product questions.We’re already putting in motion additional research projects that will expand on this initial work. We’re going to run a similar RCT for the new Figma design agent, with a specific interest in how parallel agent capabilities can further amplify time savings. We also aim to explore new hypotheses focused on AI’s impact on design quality, how it unlocks multiplayer collaboration, and beyond.