Data visualization and analysis
Publisher’s description
This introduction to visualization techniques and statistical models for second language research focuses on three types of data (continuous, binary, and scalar), helping readers to understand regression models fully and to apply them in their work. Garcia offers advanced coverage of Bayesian analysis, simulated data, exercises, implementable script code, and practical guidance on the latest R software packages.
The book, also demonstrating the benefits to the L2 field of this type of statistical work, is a resource for graduate students and researchers in second language acquisition, applied linguistics, and corpus linguistics who are interested in quantitative data analysis.
Access book files files at http://osf.io/hpt4g.
Highlights
- Intro to R
- Focus on data visualization
- Linear, logistic, and ordinal regression
- Hierarchical (mixed-effects) models
- Chapter on Bayesian data analysis
- Comprehensive code that can be fully reproduced by the reader
- File organization with RProjects
Highly recommended as an accessible introduction to the use of R for analysis of second language data. Readers will come away with an understanding of why and how to use statistical models and data visualization techniques in their research.
Lydia White, James McGill Professor Emeritus, McGill University
Curious where the field’s quantitative methods are headed? The answer is in your hands right now! Whether we knew it or not, this is the book that many of us have been waiting for. From scatter plots to standard errors and from beta values to Bayes theorem, Garcia provides us with all the tools we need—both conceptual and practical—to statistically and visually model the complexities of L2 development.
Luke Plonsky, Professor, Northern Arizona University
This volume is a timely and must-have addition to any quantitative SLA researcher’s data analysis arsenal, whether you are downloading R for the first time or a seasoned user ready to dive into Bayesian analysis. Guilherme Garcia’s accessible, conversational writing style and uncanny ability to provide answers to questions right as you’re about to ask them will give new users the confidence to make the move to R and will serve as an invaluable resource for students and instructors alike for years to come.
Jennifer Cabrelli, Professor, University of Illinois at Chicago
[…] this book’s strength lies in giving readers just enough to enable them to quickly apply their newly acquired knowledge and skills to their own data in order to produce complex, journalworthy analyses. The book is timely, with increasing expectations for more refined accounts of the diverse populations and intricate results stemming from studies of second language acquisition and bi/plurilingualism, as well as other fields of linguistic research.
Senécal & Sabourin (2023)
News & updates
Here are some updates and additional info related to the code used in the book. Some of these are based on questions I get about this code. This page will change from time to time to reflect updates in relevant packages and functions used in the book (e.g., mutate_...(); see here).
- The function
mutate_if()has been superseded byacross().
- Before:
mutate_if(is.character, as.factor) - Now:
mutate(across(where(is_character), as_factor))
- Besides using
scale_x_discrete(label = abbreviate)to abbreviate axis labels, you can also usescale_x_discrete(labels = c(...)), which allows you to choose how labels are abbreviated. - For guidelines regardings Bayesian analyses, see Kruschke’s recent paper Bayesian Analysis Reporting Guidelines.
- With R 4.1+, the native pipe
|>can replace%>%(read more here). -
guide = FALSEis now deprecated and should be replaced byguide = "none". - Instead of
select(vars), wherevarsis a vector containing multiple columns of interest, you should now useselect(all_of(vars)). - Check out the useful changes to vector functions in
dplyr1.1.0 here - You can use
read_csv()andbind_rows()(together withlist.files()andfull.names = T) to combine multiplecsvfiles in a directory (no for-loop is needed)—thanks to Natália B. Guzzo for pointing this out. - When you’re working with factors, the function
fct_relevel()from theforcatspackage offers a lot more flexibility than the functionrelevel().
Useful packages and functions not mentioned in the book
- The
dtplyrpackage provides the power ofdata.tablewith the familiartidyversesyntax. Check it out here. -
case_when()(fromdplyr) is a great function to avoid using multipleif_else()s. See documentation here. -
sample_n()will print four (by default) random rows of your data. -
dplyr 1.1.0+now offers pre-operation grouping with the.byargument within functions such asmutate()andsummarize(). A key advantage is that we no longer need toungroup()variables after applying the function. Check out my blog posts on pre-operation grouping and on snippets to automate your tasks.
- R
4.2.0and4.3.0also provide some neat features, such as the use of_as a placeholder for the native pipe|>(basically, the equivalent to.when you use%>%). In addition, you can extract specific values that are output in a pipeline in a clean and elegant way. For example, the code below extracts coefficients of a linear model.
data |> lm(response ~ predictor, data = _) |> _$coef- Typo on p. 152, paragraph 2: “we already know the probability”.
- Typo on p. 218. For some mysterious reason, the published version of the book has “Monte Carlos Markov Chain” (first paragraph), which should obviously read “Markov Chain Monte Carlo”. This is, in fact, what is listed in the glossary at the end of the book on p. 254.
- Clarification on p. 225, paragraph 2 (interpreting code block 55): “First, we have our estimates and their respective standard errors”. Bear in mind that in Bayesian models, the standard error of the coefficients corresponds to the standard deviation of the posterior distribution.
- Clarification regarding the function
mean_cl_bootmentioned on p. 95, paragraph 1: this function (from theHmiscpackage) computes bootstrapped confidence intervals. It resamples the data with replacement (1,000 times by default), computes the mean of each resample, and uses the 2.5th and 97.5th percentiles of those means as the 95% interval. Thus, the resulting error bars represent confidence intervals, not standard errors, and at 95% they will be wider than error bars representing standard errors (roughly ±2 SE). For intervals based on the normal approximation, usemean_cl_normalinstead. - On p. 20, “Simply go to RStudio > Preferences (or hit Cmd + , on a Mac)”. In more recent versions of RStudio, this has changed to “Tools > Global Options…” (same shortcut as before).
- Chapter 10 (Going Bayesian): Note that the output of a
brm()model shows the two-sided 95% credible intervals (l-95%CI andu-95%CI) based on quantiles. If the posterior is symmetrical (i.e., approximately normal), this interval will practically coincide with the highest density interval, HDI (this is the case in the chapter). However, in asymmetrical distributions, the interval shown in the outputbrm()will not coincide with HDIs.
Chapter 7 (Logistic regression)
- Typo on p. 145, paragraph 3: “while a logistic regression predicts” (not “variable”).
- p. 146, paragraph 3: “ln(0.25) = −0.60, and ln(4) = +0.60” should read “ln(0.25) = −1.39, and ln(4) = +1.39” (the values printed are base-10 logarithms). This matches Table 7.1 (P = 0.80 → odds = 4 → ln(odds) = 1.39).
- p. 147, Figure 7.3: for the same reason, the two bullets on the log-odds line should be at ±1.39 (between 1 and 2 on each side), not at ±0.6. The odds line is correct, and the point of the figure (log-odds are symmetrical) is unaffected.
- p. 148, §7.1.1: “the concept of residuals doesn’t apply” is too strong. Logistic models are not fit by least squares, and residuals cannot be computed on the log-odds scale (the observed 0s and 1s would be at ±∞), but residuals do exist on the probability scale (e.g., \(y - \hat{p}\)), as well as deviance and Pearson residuals. They are used later in the chapter: the residual deviance (p. 153) and the binned residual plot (p. 170). What does not carry over from linear regression is the least-squares fit and the classical \(R^2\).
- pp. 152, 155, 158: the notation \(e^{|\hat\beta|}\) is used to express “by a factor of” in the direction of the effect. The general conversion from log-odds to odds is \(e^{\hat\beta}\): for \(\hat\beta = -0.67\) (p. 158), \(e^{-0.67} = 0.51\), that is, the odds are multiplied by 0.51 (equivalently, divided by 1.95).
- p. 152, last paragraph: “a positive difference of 30%” means 30 percentage points (from 0.20 to 0.50).
- p. 157, paragraph 1: “\(H_0: \hat\beta_0 = 0\)” should read “\(H_0: \beta_0 = 0\)”: hypotheses are about parameters, not estimates.
- p. 157, paragraph 2: a negative intercept does not mean that the NoBreak condition “lowers” the probability of choosing low attachment. The intercept simply is the log-odds for NoBreak, which is below 0 (i.e., \(P < 0.5\)).
- Typo on p. 157, paragraph 3: “In a continuous predictor” (not “factor”).
- p. 161, paragraph 1: zero log-odds corresponds to a 50% probability only for the intercept. For all other estimates, \(H_0: \beta = 0\) means no change in log-odds relative to the reference level (i.e., an odds ratio of \(e^0 = 1\)), regardless of the baseline probability.
- pp. 164–165: the estimates for
ConditionHigh,ConditionLow,ProficiencyInt, andProficiencyAdvare not “the predicted log-odds” in each condition/group; they are the change in log-odds relative to the intercept (NoBreak, Nat). For example, the predicted log-odds for native speakers in the High condition is \(\hat\beta_0 + \hat\beta_{High} = 0.53 - 1.02 = -0.49\) (\(P \approx 0.38\)). - pp. 165, 167, 169: code block 39 and Table 7.2 list
ProficiencyIntbeforeProficiencyAdv, but running the book’s code (where Nat is set as the reference after alphabetical ordering) listsProficiencyAdvfirst, as in code block 38. Code block 40 extracts estimates by position and assumes the latter order. A safer approach is to extract them by name, e.g.,coef(fit_glm4)[["ProficiencyAdv"]]. Figure 7.7 is correct. - p. 166, paragraph 1: “\(\hat\beta = 1.08\)” should read “\(\hat\beta = 1.07\)” (see code block 39 and Table 7.2: 1.072).
- Typo on p. 166, Reporting results: “significantly affected by” (not “affect”).
- pp. 167–168: the text refers to the levels of
Sigas “Yes” and “No”, but code block 40 uses lowercase"yes"and"no". - p. 168, paragraph 2: Wald tests do not compare a model with and without predictor \(X\); that describes the likelihood ratio test (
anova(..., test = "LRT")). A Wald test divides each estimate by its standard error (\(z = \hat\beta/SE\)) and compares the result to a standard normal distribution: this is thez valuein the output ofsummary(). - p. 170, paragraph 1: the binned residual plot should use fitted probabilities and raw (response) residuals:
binnedplot(fitted(fit_glm4), residuals(fit_glm4, type = "response")). With the defaults shown in the book,predict()returns log-odds andresiduals()returns deviance residuals, which do not match the confidence bands drawn bybinnedplot().
How to cite
Garcia, G. D. (2021). Data visualization and analysis in second language research. New York, NY: Routledge.
@book{garcia_2021_dvaslr,
title = {Data visualization and analysis in second language research},
author = {Garcia, Guilherme Duarte},
year = {2021},
address = {New York, NY},
publisher = {Routledge},
isbn={9780367469610}}
Part of this project benefited from an ASPiRE Junior Faculty Award at Ball State University (2020–2021).
Copyright © Guilherme Duarte Garcia