Factor analysis allows us to understand more about scales and measures used in psychology research. At its core, factor analysis asks: "What hidden variables might explain why these observable (or measurable) variables tend to go together?" It's a way of finding unobservable structure in a set of variables. When many measured items consistently rise and fall together, factor analysis proposes that they're being driven by an underlying, unobserved construct (a factor).
The APA Dictionary, with more than 25,000 entries, is a good place to start when you spot a statistical analysis new to you. The dictionary encompasses lots of statistical terms, including factor analysis and many terms related to factor analysis, which can help you to start to understand what’s going on.
APA defines factor analysis as “a broad family of mathematical procedures for reducing a set of interrelations among manifest variables to a smaller set of unobserved latent variables or factors.” Let’s break that down. It’s a broad family of procedures, which means it’s not just one thing. In fact, we won’t be able to cover all of the variations in this brief module. The word interrelations is important because it brings up the idea of relationships among variables. If this makes you think of the central statistical concept of “correlation,” you’re already on track to understanding factor analysis! More on that later.
Finally, you remember variables as the basic building blocks of statistical analyses. In factor analysis, manifest variables are the ones you can “see” – that is, you can observe or measure them. Latent variables are the hidden variables; we cannot observe or measure them, so we must approximate (or estimate) them by analyzing manifest variables. Factor analysis is an important tool in identifying and understanding latent variables!
To pull it all together, factor analysis, broadly defined, looks for relationships among measured variables to understand an underlying, or hidden, structure among those variables.
Let’s consider an example. The Beck Depression Inventory, or BDI for short, is a self-report measure that includes 21 items. For example, one BDI item asks respondents to provide a score from 0 (“I don’t get tired more than usual”) to 3 (“I am too tired to do anything”). Scores on all 21 items are added up to provide an overall depression score. The individual items on the BDI, as well as the overall BDI score, are manifest variables. We measure them by asking people to respond to the items on the BDI. But are all 21 items measuring the same thing?
One factor analysis was conducted on BDI data from several hundred Portuguese people, some of whom had cancer (Almeida et al., 2023). The factor analysis statistically identified three underlying latent variables – three categories of symptoms assessed by the 21 items on the BDI. These are
cognitive symptoms (items related to thinking) like blaming yourself when bad things happen or having difficulty making decisions;
affective symptoms (those related to moods or feelings) like feeling sad or pessimistic; and
somatic symptoms (those related to bodily symptoms) like changes in appetite or sleeping habits.
These three are latent variables. (Important caveat: There are different ways to do factor analysis, and lots of different samples of participants in which you could examine the BDI. Because of that, not all researchers have found three latent variables or have categorized the 21 symptoms in exactly the same ways, although there is a lot of overlap.)
Here are just a few possible research questions, which we chose because they span disciplines:
Psychology: Do 20 items on a stress questionnaire really measure one thing, or are there distinct sub-types (e.g., emotional stress vs. physical stress vs. social stress)?
Marketing: What underlying dimensions drive consumer satisfaction with a product? (e.g., "quality," "value," "aesthetics")
Public health: Do neighborhood characteristics (walkability, green space, air quality, noise, access to food, and so on) cluster into a smaller number of "neighborhood quality" factors?
Political science: Do voters' positions on many different policy issues reduce to a few underlying ideological dimensions?
Education: Does a 30-item intelligence test actually measure several distinct abilities (verbal, spatial, memory), or one general factor?
How to interpret in-text results: Let’s walk through an example using an open-access article. Before continuing, go ahead and download the article, Confirmatory factor analysis and exploratory structural equation modelling of the factor structure of the Depression Anxiety and Stress Scales-21, by Rapson Gomez, Vasileios Stavropoulos, and Mark D. Griffiths: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0233998. We’ll first give some generic language to look for, and then talk about how the Gomez et al. (2020) uses that language. As you read, find the language we discuss in the article.
THE ANALYSIS AND GENERAL FINDINGS: Most factor analysis papers will say something like: "A principal components analysis / exploratory factor analysis was conducted...": There are two main kinds of factor analyses, exploratory and confirmatory. In exploratory factor analysis (EFA), the researchers let the data reveal the structure. Gomez et al. use confirmatory factor analysis (CFA), but they clearly explain the difference early on, noting that with EFA, “items are enabled to associate with various dimensions concurrently.” Results might say something like “three factors were extracted.” CFA assumes a structure in advance, usually based on theory or previous research. As Gomez et al. explain, CFA “allows the researcher to test for a priori defined factor structure.” (A priori essentially means decided in advance.) Results might say something like “the three-factor model showed good fit.” Either way, factor analysis is asking whether a set of items clusters into a smaller number of meaningful groups.
As for the Gomez et al. article, the Depression Anxiety and Stress Scales-21 (DASS-21) was designed to measure three things: depression, anxiety, and stress. Gomez et al. want to test it to be sure this breakdown is reflected in the data. The question isn't "how many factors are there?" but "does the three-factor structure we predicted fit the data?" The authors tested this model in a few different ways, and report “For all models, their CFI, TLI and RMSEA values indicated good fit.” They provided the numbers in Table 2 (CFI = .969, TLI = .965, RMSEA = .049) but you can just read their description of the findings, rather than these numbers, to understand whether the predicted factor structure emerged.
[Digging deeper: In case you’re curious, CFI is comparative fit index, TLI is Tucker-Lewis Index, and RMSEA is root mean square error of approximation. CFI and TLI range from 0 to 1, and higher is better – indicating good fit. RMSEA measures error, so lower is better; generally anything below .06 is good.]
FACTOR LOADINGS: You will see language similar to this: "Factor loadings ranged from .45 to .78." A loading tells you the degree to which an item “belongs” to a factor. You can think of factor loadings as reflecting a correlation between the item and the latent (invisible) factor. Here are some examples from Table 3 in the Gomez article (we’re using just one of their models here, the CFA three-factor model in the first several columns):
“Life was meaningless” loaded .81 on the depression factor, unsurprising given that this item is a key marker of depression.
“Felt close to panic” loaded .81 on the anxiety factor, also unsurprising.
“Working up initiative loaded .65 on the depression factor, still solidly on the factor but a weaker marker
“Dryness of mouth” loaded .40 on anxiety, which is on the lower end.
Gomez et al. use a word that you will often see in factor analysis articles – salient. In everyday language, something is salient if it is prominent, striking, or noticeable. In factor analysis, it means that a factor loading indicates a strong, meaningful relationship with the underlying variable. Gomez et al. treat a factor loading as salient if it is at least .30.
It can be really interesting to explore a table of factor loadings and see which items are more strongly correlated with a given factor and which are less so. It can give you insights into a specific construct, such as anxiety.
VARIABILITY AND FACTOR ANALYSIS: In many factor analyses, particularly exploratory ones, you’ll see language such as "Three factors were extracted, accounting for 62% of the variance." You probably remember from intro to stats that variance is the spread in people’s scores. In factor analysis, variance “accounted for” asks how much of that spread the factor structure explains.
Gomez et al. never use this language, however, because they conducted a confirmatory analysis. Instead, they refer to fit, a different way to refer to variability. Gomez et al. use a related idea, however: ECV (explained common variance). The depression, anxiety, and stress factors had ECVs of .43, .30, and .26, which tells us that of the variance the items share, depression pulls the biggest part.
Generally, EFAs talk about variance explained and CFAs talk about fit.
ROTATION: In most factor analyses, you’ll see a discussion of one or more kinds of rotation. Rotation is a mathematical technique for making the factor structure cleaner and more interpretable, almost like adjusting the angle of a lens until a picture comes into focus. You might see varimax (an orthogonal rotation in which factors are forced to be uncorrelated) or promax (an oblique rotation in which factors are allowed to correlate). In the Gomez et al. article, we see target rotation which aims the solution at a predicted pattern. And Gomez et al. used oblique rotation, allowing their factors to correlate, which they did. The correlations among the three factors (depression, anxiety, stress) were high, ranging from .73 to .90 in one of the tested models (see Table 2).
TABLES AND FIGURES IN FACTOR ANALYSIS: There are several common tables and factors you’re likely to encounter as you read a factor-analysis article.
Factor loading table: This is Table 3 in the Gomez et al. article. The 21 items are in the first column, and additional columns show the factor loadings for the different models the authors tested. See part 2 of this module for more information on factor loadings.
The fit table: This is Table 2 in Gomez et al. You can see one row per model tested with several statistics represented in the columns under “Fit Values”: χ² , CFI, TLI, and RMSEA. (The factor correlations are in the remaining columns.)
Path diagrams: See Figure 1 in Gomez et al. The circles are the latent (invisible) factors, with D, A, and S referring to depression, anxiety, and stress, respectively. The boxes are the items, what we’re actually measuring. The arrows represent factor loadings. Let’s look at the path diagrams for each of the four models. For the one titled “ICM CFM 3-factor,” each box gets one arrow, indicating that each item loads on to one, and only one, factor. The ESEM 3-factor one looks like a mass of tangles, indicating “cross-loadings” showing items correlating with non-target factors, too. The two lower diagrams add a big G circle to the left, a general distress factor that all items feed.
Scree plots: When you read about EFA, you’ll encounter scree plots. Because Gomez et al., is a CFA, they do not use a scree plot. So, let’s look at this hypothetical one. Each blue circle shows how much variance one component (or factor) explains, what is called its eigenvalue. The line starts high with the first component explaining a lot (eigenvalue around 4). It then plunges. The second component explains about 2, the third 1, and then the dots trail off in a nearly flat line. That flat stretch is the “scree,” named for the rubble at the bottom of a steep mountain slope. The bend where the steep becomes flat is called the “elbow.” Components before the elbow are considered worth keeping; those after are “noise.” Another criterion is an eigenvalue of 1. Anything above it stays. Here, two factors are clearly above 1, and one factor is right at it, so this indicates either a two- or three-factor solution. It’s important to note that there is subjectivity to the scree plot, and the researcher must weigh in, using theory and data, in this case to decide between two and three factors. Moreover, not all researchers believe they are a rigorous or helpful tool. Nonetheless, you’re likely to see them.
MEASURES OF INTERNAL CONSISTENCY: There are several measures of internal consistency, including Cronbach's alpha and omega (ω). Such measures indicate how strongly a set of items hangs together, on a 0-to-1 scale, with higher indicating more consistency. They are often reported alongside factor analysis, and both alpha/omega and factor loadings are ways of asking "do these items belong together?" Gomez et al. use omega and report internal reliabilities of .84 (depression), .71 (anxiety), and .83 (stress) for their sample. There are different criteria for what is considered good internal consistency, so check the article for what the authors used. Gomez et al. report that they used a criterion whereby “ω values need to be at least .50 with values of at least .75 being preferable.”
How to interpret in-table results: We’ve already discussed factor loading tables, which are the main artifact of a factor analysis. We’ll give you a quick overview of interpreting these tables here, using Gomez et al. as the example when possible.
Rows = individual items (For Table 3, this includes all 21 DASS items, by brief description such as “felt down-hearted” or “trembling”; Columns = factors (D, A, S, plus G, the general factor, for the bi-factor models). Remember that Gomez et al. conducted factor analyses using four different models, all of which are represented in this table. Factor loadings are at the intersections of rows and columns for each of the four models. Note that the empty cells in the CFA model are not missing data; rather, they are loadings fixed to zero by that particular model.
Larger numbers (closer to ±1) = stronger item-factor relationship. This is the same scale and same interpretation as correlation. “Life was meaningless” in the ESEM depression column is practically the definition of depression. It’s important to note that loadings, like correlations, have direction, and there can be negative loadings. For example, “Hard to wind down” loaded -.33 on the anxiety factor for the ESEM model. A negative loading on a targeted factor is a red flag, one that the authors explicitly discuss (search “negative loadings” in the article to see what they say),
Cross-loadings (an item loading on two factors) = interpretation complexity. In the ESEM model, the stress item “Tended to over-react” loaded on stress (.32) but also on depression (.14) and anxiety (.36). And the anxiety factor “Dryness of mouth” loaded more on stress (.58) than its own factor (-.14). When items lean on multiple factors like this, the factors are blurry, which is why these authors concluded that some factors were “poorly defined.”
Asterisks and significance: The asterisks in the table indicate statistical significance of individual loadings. You can look below the table to see what asterisks represent. For Gomez et al.’s Table 3, one asterisk indicates p < .05, and two indicates p < .01.
Defining a factor: Items that define a factor are usually those with loadings ≥ .30 or .40. (Gomez et al. used .30.). Remember that the term for an item over the chosen threshold is “salient.” There are many, many ways to define a factor. If in doubt, read what the authors of the article you’re reading decided, and why.
Alpha/omega in tables: When reliability coefficients appear in a table, they describe each factor, with all of its items, as a whole.
How to interpret a data visualization: We discussed the basics of understanding path diagrams and scree plots in part 1e. Here we want to ask you to dig deeper. Look at Figure 1 in Gomez et al. Compare the CFA model (upper left; clean, one arrow per item) with the ESEM model (upper right, messy, arrows everywhere). Which diagram do you think better matches psychological reality? We’re guessing most of you will say the messy ESEM model. Yet, the authors chose the cleaner CFA model anyway (we’ll discuss more in part 4 below)! Path diagrams illuminate in ways that can lead to thoughtful discussion and debate.