Teaching Writing with Quantitative Data
Students in most fields will need to write with numbers, but the ways that quantitative data are used vary across courses and curricula. In experimental and scientific courses, students may be taking measurements and recording outcomes of tests. In social science courses, students will interpret statistical data to describe populations and to draw inferences. Even in arts and humanities courses, numbers pervade disciplinary processes of description and analysis.
Although the old adage suggests that “numbers speak for themselves,” the ways in which we present quantitative data and the ways we select that data for our purposes and audience make a huge difference in the success of documents. Unfortunately, students often learn the algorithmic processes of deriving, reporting, and analyzing numbers without considering the contexts, purposes, and audiences that turn numbers into insight. Furthermore, with the abundance of GenAI tools, students are now able to generate quantitative data from black-box processes that do not involve their input or calculation. They may be tempted to cognitively offload the reading and writing with numbers to Gemini and other AI tools—tools that can include hallucinations, misapplied statistical tests, or incorrectly applied mathematical formulas. Providing students the opportunity to write with and about data will allow them to use and develop their computational and analytic skills to inform and persuade. On this page, we identify strategies for helping students develop effective habits for writing with data.
Preparing Your Students to Write with Data
Numerous statisticians, economists, mathematicians, and psychologists have independently reached the conclusion that humans are relatively weak intuitive statisticians. We have a very difficult time understanding probability in context, especially when we consider risks or consider outcomes in which we have an interest. We are often very poor at understanding the differences between normal variation, biases, and statistical noise. We suffer from recency effects, halos, and horns. Perhaps worst of all, we will often abandon the successful and accurate conclusions of statistical data if doing so allows us to retain our previously held beliefs.
Despite their processing and computational speeds, Gen AI tools are also prone to weaknesses. While GenAI tools can be used for recognizing patterns quickly in the data, summarizing statistical results, and correcting coding errors, they lack real-world, human context and can reproduce the biases that exist in the data or inputs they have trained on. Even when students are encouraged to use Gen AI to write with data, they will need the capacity to evaluate the output. Although students who write with code can no doubt find functional (if inelegant) shortcuts through GenAI, only the most inexperienced or foolish analyst would trust their statistical results to “vibe stats.”
Consider an Ungraded Pretest Activity
If your course requires a familiarity with statistics, assess your students’ statistical knowledge early in the semester. A low-stakes quiz or an informal writing activity on relevant statistical terms and concepts will help you to understand the extent of students’ background knowledge. If you find that students lack understanding of some foundational concepts, you might either consider a brief review in class or recommend supplemental instruction on statistics online. You can also model for students ways to use GenAI in conjunction with a review protocol, known as the four-step Feynman technique, developed by the Nobel award-winning physicist Richard Feynman. For example, a student can use a prompt like this to practice statistical understanding with Gemini: "I am going to explain the concept of 'normal distribution' and 'variance' in my own words. I’m going to explain it in a language that a 12-year-old would understand. Read my explanation and tell me if my logic is sound, and point out any misconceptions I might have." Providing students with pathways to practice and develop their statistical knowledge, especially early in the semester, can help to mitigate misunderstandings of statistical models and tests that may only appear in final graded work when it is too late to revise.
Frequent, low-stakes opportunities to test students’ familiarity with concepts and methods offer multiple advantages to students, for both demonstrating and building confidence. In Make it Stick: The Science of Successful Learning, Brown, Roediger, and McDaniel emphasize the value of self-testing for retrieval practice:
If instructors want to conduct ungraded pretest activities that reflect what students themselves currently understand, without AI assistance, they should do so in class, on paper, or in a locked-down quiz tool.
If instructors allow the use of GenAI tools on ungraded pretest activities, they should model for students how a GenAI tool, such as UMN’s Gemini enterprise account, can be used to generate practice quizzes with prompts such as this one, recommended specifically by Gemini: "I am studying introductory statistics. Generate a 5-question multiple-choice quiz testing the difference between standard deviation and standard error. Do not give me the answers until I respond." See Bowen and Wilson’s Teaching with AI: A Practical Guide to a New Era of Human Learning (2025) for additional tips on integrating AI into assessment.
Help Students Understand That Computational Familiarity Isn’t the Same as Selection and Application
Students may be familiar with mathematical and statistical concepts from their previous coursework, but they may not have a sense of how or why particular measures or numbers may be relevant in context. For example, most students know the difference between mean, median, and mode, but they may not understand when a median statistic is more representative than a mean, or why a mode may be of interest with a particular data set. Students may be able to describe a correlation mathematically, but they may not know whether a correlation is meaningful or if an effect size is large enough to matter. Providing students with cases and context-rich opportunities for application can help them keep the big picture in mind.
This distinction matters even more in Gen AI-assisted writing environments. Gen AI tools can define concepts like mean, median, and mode, or describe what a correlation coefficient measures. However, they are less reliable at judging whether a median is more representative than a mean for a particular data set, or whether a correlation is meaningful enough for discussion. Furthermore, Gen AI tools lack access to a course's disciplinary context, a study's real-world stakes, or an audience's specific needs. This can lead to recommendations without any of the judgments those decisions require. Instructors can model these inconsistencies in class working with a GenAI tool, to help students understand that no matter how fluently a tool can describe a statistical concept, choosing and applying the right one for a given context remains the students’ responsibility.
Be on the Lookout for Common Statistical Misconceptions (Especially Counterintuitive Conceptions)
Even though students may be familiar with positive and negative correlations, they may still misapply their knowledge of positive and negative numbers to assume that variables that rise together are positively correlated and variables that are negatively correlated fall together (rather than diverge).
Perhaps the most common statistical errors emerge when statistical terms with technical meaning are misunderstood by their conventional definitions. For example, students may confuse significance (mathematical evidence of a non-random effect) with significance (something meaningful or important). Students may inadvertently associate positive correlations with preferred outcomes and negative correlations as bad news. In their analysis of specialized vocabulary in statistics, Kaplan, Fisher, and Rogness noted that ambiguity around key statistical terms (‘average’, ‘confidence’, ‘random’, and ‘spread’) created challenges for students in their early statistics learning. Providing students with opportunities to write with these terms in their specialized statistical context can help them master important statistical concepts. Likewise, Gen AI tools may use common vernacular terms rather than precise statistical terminology. Instructors should be clear with students that if they use GenAI to help draft text, they must carefully review the AI output to ensure terms like significance, variance, or correlation are being used in their strict statistical sense, not colloquially.
Selecting the Best Tools: Describing and Representing Data
Although quantitative reasoning is common across the curriculum, the means by which disciplines gather, report, synthesize, and attach significance to data can differ dramatically. In the previous section, we looked at ways to use writing to address potential conceptual problems, in this section, we emphasize teaching disciplinary practices of reporting and representing data.
Identify the Common Reporting Practices of Your Field
When addressing course readings or looking at examples, it can be valuable to reinforce which measures and tests are included and why those measures are considered important to the field. In addition to describing the quantitative techniques involved, students will benefit from descriptions of how and why the technique involved applies to the context. Ideally, students will not only build technical skills with the tools of the field, but they will also have a system for organizing and selecting the appropriate tool for their data tasks. Further, identifying the “why” of a particular mathematical or statistical technique will help students to distinguish between merely describing a process and using that process to generate conclusions. Using pedagogical techniques like metateaching or even merely being more explicit about why certain reporting practices are the best tools can help students to understand the sometimes implicit values of disciplines.
Provide Context-Rich Problems and Questions
Conversely, when students are turned over to their own quantitative analyses, providing opportunities to draw meaningful conclusions will reinforce how and why a particular quantitative technique is useful or applicable. In many fields, small case studies from professional or laboratory contexts can give students more opportunities to connect their quantitative analyses to meaningful conclusions. Similarly, scaffolding larger assignments can provide opportunities for meaningful student learning and effective feedback on writing.
Insist on Consistent, Careful Attention to Units of Measurement
A common error students make in their writing with quantitative data is omitting units of measure. When students are preparing writing for assessment by an instructor or teaching assistant, they may (correctly) assume that their reader already knows the context, problem, and answer and is merely interested in “the number” that provides the answer. Asking students to consistently include units of measure (even with problem sets and homework) reminds them that in most real-world instances, analyses are provided to audiences who may need to be supplied with elements of context in order to draw appropriate conclusions.
Ask Students to Express Their Answers in a Sentence
One of the simplest strategies for assisting students in writing about numbers is to give them practice. Requiring students to write their answers in a complete, meaningful sentence can be a valuable habit. For example:
The best answers to questions requiring data should include relevant context (who, what, when, where) and ascribe significance (why and how it matters).
Turning a number into a meaningful, audience-aware sentence is a skill tempting to hand off to a GenAI tool, but it is a skill students should practice themselves. To ensure that students are doing the learning, consider doing this in class as a scaffolded writing-to-learn activity so students can practice and build the habit themselves before deciding whether and how AI might assist.
Data Analysis and Synthesis: Drawing Conclusions to Inform and Persuade
The terms “data analysis” and “data synthesis” are often used interchangeably, but they might be usefully distinguished by the purposes for which data are collected and used. Data analysis typically involves the application of statistical tests or formulae to draw conclusions based on data. Most efforts to draw generalizations, conclusions, or predictions involve computational work. Data synthesis occurs when someone brings together data from multiple sources for the purpose of aggregation or presentation. Systematic literature reviews or observational studies will often involve data synthesis. Again, our common understanding of these terms can be confusing, especially if one of these terms is used simply to mean “make meaning from numbers.”
First, Help Your Students Differentiate Data Collection from Data Analysis
When presenting their data, students sometimes merely describe their processes for collecting data in a sequence of steps rather than providing analysis and synthesis of results. Students may misunderstand the process of describing methods to mean offering an account of what they did, rather than as a description of methodology and its value and limitations. If students' data collection is to be assessed, be clear about why raw data is necessary. If only students' conclusions are most important, remind them that displaying raw data may not be necessary. Providing students with annotated examples can help them to recognize which information is relevant to different audiences.
Identify Examples of Successful Data Analysis in Course Readings and Literature
When students interact with their course readings, they may be tempted to focus on the what—what is described?—rather than the how--how do authors use evidence to make their case? By connecting the organization of written material in your course to the purposes and audiences they serve, instructors can advise students to keep audience needs and genre expectations in mind. Simple, authentic examples that connect claims to data can be especially useful in helping students understand synthesis. For especially complex readings or foundational texts, the use of tools like social annotation can assist students in recognizing the features of successful writing with data.
Provide Systematic reviews and Meta-Analyses to Illustrate Data Synthesis
Some research genres depend primarily on aggregating previously recorded data to explain the state of knowledge on a topic or to illustrate broader claims about previous research. These explicitly synthetic forms of writing can be meaningfully contrasted with other types of quantitative research, whether they involve collecting data or querying existing data sets. Some disciplines may even be quite explicit about the relative value of different methods of data collection for supporting evidence-based practice.
Synthesis assignments do carry a GenAI risk. GenAI tools are prone to hallucinations, such as fabricating citations, misreporting effect sizes, or attributing findings to studies that said something different (or do not exist at all). Any assignment that asks students to synthesize prior research, with or without AI assistance, should require students to verify every cited statistic and claim against the primary source themselves before it appears in their writing.