Proteomics

MaxQuant 2.8 proteinGroups.txt, Scored Against a Known Answer: LFQ vs iBAQ vs Intensity, and What Match-Between-Runs Adds

I ran MaxQuant 2.8.1 on six public Orbitrap runs where human, yeast and E. coli were mixed at known ratios, then scored every intensity column against the right answer. LFQ was far more accurate than raw Intensity but quantified fewer proteins, iBAQ ratios were identical to Intensity, MBR made already-quantified proteins more accurate, and the old Reverse filter silently returns zero rows in 2.8.

·12 min read
#MaxQuant#MaxQuant 2.8#proteinGroups.txt#LFQ intensity#iBAQ#match between runs#label-free quantification#benchmark#proteomics R#downstream analysis

MaxQuant proteinGroups scored against known ratios

Why Score MaxQuant Against a Known Answer

Most MaxQuant tutorials stop when proteinGroups.txt appears. The questions that matter come after: which of the three intensity columns to use, what the filter columns actually remove, and whether match-between-runs (MBR) adds real quantities or noise.

On real samples you cannot answer those, because nobody knows the true fold changes. So I used a public benchmark where the answer is known, ran MaxQuant 2.8.1 on it, and scored every column.

This is the MaxQuant counterpart to the DIA-NN 2.x downstream post. If you are setting up a run, the MaxQuant tutorial and the pre-flight checklist cover the part before this one.

The Dataset: Three Species at Known Ratios

PRIDE PXD028735 (Van Puyvelde et al., Scientific Data 2022) mixes human, yeast and E. coli digests in two fixed proportions:

HumanYeastE. coli
Condition A65%30%5%
Condition B65%15%20%
Expected log2(B/A)0−1+2

I used the Orbitrap Q Exactive HF-X DDA runs (120-minute gradient). The paper made each condition as three separate master batches. The file names carry Alpha, Beta and Gamma, which I take to be those batches, plus numbered repeat injections of each. I took one injection per batch (3 vs 3, 6 runs, 21.7 GB) so that repeat injections are not counted as replicates (the same trap as 210 runs that were really n = 70).

The Run

SoftwareMaxQuant 2.8.1.0, MaxQuantCmd on Linux (WSL2 Ubuntu 24.04), .NET 8 runtime
FASTAUniProt Swiss-Prot: human (20,431), yeast S288c (6,733), E. coli K-12 (4,531) + MaxQuant contaminants
SettingsTrypsin/P, 2 missed cleavages, Carbamidomethyl (C) fixed, Oxidation (M) + Acetyl (protein N-term) variable, 1% PSM and protein FDR, LFQ (min. ratio count 2), MBR on, iBAQ on
Machine16 of 32 threads, 62 GB RAM available to WSL
Wall time61.5 minutes for 6 runs · peak memory of the main process ~11 GB

The slowest steps were retention-time alignment (9.2 min, part of MBR), the first search (6.7), mass recalibration (5.7) and the second-peptide search (5.6). Each run had 132,600–138,100 MS/MS spectra submitted, 29–30% of them identified, and 27,600–28,700 peptide sequences; 42,760 sequences across all six.

A first attempt with 28 threads crashed the whole WSL instance partway through; the most likely cause was the Windows drive that holds WSL's virtual disk running out of space. If you run MaxQuant under WSL, watch the free space on that Windows drive as well as what df reports inside Linux. The virtual disk also has its own size cap.

Gotcha 1: --create truncated my experiment names

MaxQuantCmd --create writes a parameter file from a raw folder. It named the experiment for LFQ_Orbitrap_DDA_Condition_A_Sample_Alpha_01.raw as A_Sample_Alph: prefix dropped, last letter cut. LFQ groups and labels everything by experiment name, so check the <experiments> block before running and set clean names yourself:

<experiments>
   <string>A_Alpha</string>
   <string>A_Beta</string>
   ...
</experiments>

Gotcha 2: There Is No Reverse Column Any More

In the 2.8.1 output, the decoy flag column is called Decoy. Scripts written for older versions usually filter on Reverse. In base R that filter does not fail. It silently empties your table:

pg <- read.delim("proteinGroups.txt", check.names = FALSE, quote = "")
nrow(pg)                         # 5785
nrow(pg[pg$Reverse != "+", ])    # 0   <- no error, no warning
subset(pg, Reverse != "+")       # Error: object 'Reverse' not found

pg$Reverse is NULL, the comparison returns a zero-length vector, and indexing with it keeps nothing. If a downstream step then reports "0 proteins" you might look everywhere except the first filter. Check the column first:

decoy_col <- intersect(c("Decoy", "Reverse"), names(pg))
stopifnot(length(decoy_col) == 1)

What the Three Filters Remove

Protein groups
Rows in proteinGroups.txt5,785
Decoy = "+"68
Potential contaminant = "+"59
Only identified by site = "+"44
After all three5,628
… human / yeast / E. coli / mixed-species3,780 / 1,457 / 389 / 2

E. coli is the smallest group, partly because its proteome is smaller and partly because it is only 5% of the material in condition A. Low-abundance proteins are the first to drop out, and that matters for every result below.

Gene names holds more than one name, separated by semicolons, in 172 groups (3.1%): the identified peptides cannot tell close homologs apart. Do not split those into separate rows (that double-counts the same intensity). Either keep the group label as it is, or report the first name and say so:

pg$gene <- sub(";.*", "", pg$`Gene names`)   # first name only; say so in your methods

Which Intensity Column? The Scores

For each protein group quantified in at least two of three runs in both conditions, I computed log2(B/A) from the mean log2 intensity and compared it with the known answer. "Within ±0.5" is the share of proteins whose ratio landed within 0.5 log2 units of the truth.

ColumnProteins scored (human / yeast / E. coli)Within ±0.5: human (0)yeast (−1)E. coli (+2)
Intensity3,603 / 1,250 / 24386%74%43%
iBAQ3,603 / 1,250 / 24386%74%43%
LFQ intensity2,896 / 907 / 13299%92%81%

LFQ is much more accurate, and quantifies fewer proteins. The spread (interquartile range) of the human ratios is 0.14 with LFQ versus 0.28 with raw Intensity; for E. coli it is 0.26 versus 0.95. But LFQ requires at least two peptide ratios per protein pair, so it leaves 20% fewer human and 46% fewer E. coli proteins quantified. Across all six runs, 74% of LFQ cells have a value versus 93.5% for Intensity. Which one you want depends on whether a missing protein or a wrong ratio hurts your study more.

iBAQ ratios are identical to Intensity ratios. iBAQ is the summed intensity divided by the number of theoretically observable peptides of that protein. The divisor is the same in every sample, so it cancels in any between-sample ratio. iBAQ is for comparing different proteins within one sample (rough absolute abundance); it adds nothing for group comparisons.

LFQ shifted the unchanged proteins. The human median under LFQ is +0.14, not 0. MaxLFQ normalisation assumes most proteins do not change between samples. This benchmark changes about a third of them on purpose (all yeast down, all E. coli up), and that is the likely reason the normalisation moved the unchanged ones. Raw Intensity, which is not normalised, put the human median at +0.04. In a real experiment with large, one-directional changes the same thing can happen without any warning, so check where your "unchanged" proteins sit before trusting small fold changes.

What Match-Between-Runs Added

MBR transfers identifications from runs where a peptide was sequenced to runs where it was only seen in MS1. peptides.txt records, per experiment, whether each value came By MS/MS or By matching.

Of 210,449 quantified peptide values, 20,830 (9.9%) came from matching. Among peptides quantified in all six runs, about 21% were complete only because of MBR. Scoring those peptides against the known ratios (centred on the human median):

Within ±0.5 of truthAll 6 values by MS/MS1–2 runs by matching3+ runs by matching
Human (0)95% (n = 15,210)89% (3,078)91% (845)
Yeast (−1)93% (3,598)87% (844)91% (252)
E. coli (+2)88% (329)77% (116)62% (26)

The matched values are centred on the right answer, so MBR is not inventing fold changes. They are noisier, though, and the noise grows where it hurts most: in low-abundance proteins that depend on several matched values. If a protein's significance rests on matched values in one condition only, look at it before you report it.

Protein level: the same run with MBR switched off

Peptide-level noise is only half the story, because MaxLFQ builds each protein ratio from many peptide ratios. So I re-ran the same six files with identical settings and MBR off. That took 52.0 minutes instead of 61.5; MBR cost about 9.5 minutes here, almost all of it retention-time alignment.

LFQ intensityProteins scored, MBR onMBR offWithin ±0.5, MBR onMBR off
Human (0)2,8962,34699%99%
Yeast (−1)90765692%95%
E. coli (+2)1327781%62%

LFQ cells with a value went from 60.5% (off) to 74.2% (on). MBR added 23% more scoreable human proteins, 38% more yeast and 71% more E. coli.

The two accuracy columns compare different protein sets, so the cleaner test is the proteins that were scoreable in both runs:

Same proteins, LFQnWithin ±0.5, MBR onMBR offIQR onIQR off
Human2,34299.5%99.1%0.130.14
Yeast65697.7%95.0%0.150.18
E. coli7794.8%62.3%0.190.47

For the same low-abundance proteins, MBR made the LFQ ratios much better: with more peptides quantified in every run, MaxLFQ has more ratios to work with. The proteins that only became scoreable because of MBR are the weaker part. Of the 860 such protein groups, 96% of the human ones landed within ±0.5, but only 78% of the yeast and 62% of the E. coli ones.

For raw Intensity the effect was even larger, because a sum over fewer detected peptides is biased: within ±0.5 dropped from 86/74/43% (on) to 77/59/20% (off) for human/yeast/E. coli.

For this dataset, then: MBR improved the proteins you would have quantified anyway, and added new ones that are less reliable than the rest. If a result rests on a protein that is only quantified with MBR on, treat it as the weakest evidence in your table.

The Code

library(data.table)
pg <- fread("proteinGroups.txt", sep = "\t", quote = "")

decoy_col <- intersect(c("Decoy", "Reverse"), names(pg)); stopifnot(length(decoy_col) == 1)
pg <- pg[!(get(decoy_col) %in% "+") &
         !(`Potential contaminant` %in% "+") &
         !(`Only identified by site` %in% "+")]

A <- c("A_Alpha", "A_Beta", "A_Gamma"); B <- c("B_Alpha", "B_Beta", "B_Gamma")
lfq <- as.matrix(pg[, paste("LFQ intensity", c(A, B)), with = FALSE])
colnames(lfq) <- c(A, B)      # columns are "LFQ intensity A_Alpha" etc.
lfq[lfq == 0] <- NA          # MaxQuant writes 0 for "not quantified"
lfq <- log2(lfq)

ok <- rowSums(!is.na(lfq[, A])) >= 2 & rowSums(!is.na(lfq[, B])) >= 2
log2fc <- rowMeans(lfq[, B], na.rm = TRUE) - rowMeans(lfq[, A], na.rm = TRUE)
summary(log2fc[ok])

Two details that bite. MaxQuant writes 0, not NA, for a missing quantity; convert before log2, or you get -Inf. And use the same valid-value rule for every column you compare; otherwise you compare different protein sets.

A Methods Paragraph You Can Adapt

Raw files were processed with MaxQuant 2.8.1.0 against UniProt Swiss-Prot [organisms] and the MaxQuant contaminant list. Trypsin/P with up to two missed cleavages; carbamidomethylation of cysteine as fixed, oxidation of methionine and protein N-terminal acetylation as variable modifications; 1% FDR at PSM and protein level. Label-free quantification used MaxLFQ with a minimum ratio count of 2[, and match-between-runs was enabled]. Protein groups flagged as decoy, potential contaminant or only identified by site were removed. LFQ intensities were log2-transformed with zero values treated as missing, and proteins were retained when quantified in at least [k] of [n] samples per group.

FAQ

Should I use LFQ intensity or Intensity for differential expression?

In this benchmark LFQ was clearly more accurate, at the cost of fewer quantified proteins. For group comparisons, LFQ is the usual choice; keep raw Intensity in mind when low-abundance coverage matters, and check where unchanged proteins sit if you expect large, one-directional changes.

Is iBAQ better for fold changes?

No. Between samples it gives exactly the same ratios as Intensity. Use iBAQ to rank proteins by abundance within a sample.

Should I turn MBR on?

In this benchmark, yes: it made LFQ ratios of the same proteins more accurate (E. coli 95% vs 62% within ±0.5) and added 23–71% more scoreable proteins per species. The added proteins were less accurate than the rest, so treat MBR-only quantifications as the weakest evidence. This is a mixture of very similar samples; MBR is generally considered riskier when samples differ a lot.

My script filters Reverse and now nothing is left.

MaxQuant 2.8.1 calls that column Decoy. See Gotcha 2.


Data: PRIDE PXD028735. Van Puyvelde B, et al. A comprehensive LFQ benchmark dataset on modern day acquisition strategies in proteomics. Sci Data 2022;9:126. doi:10.1038/s41597-022-01216-6 (CC BY 4.0). Processing and all counts in this post are my own reanalysis of the public raw files.

Software: Cox J, Mann M. MaxQuant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification. Nat Biotechnol 2008;26:1367–1372. Cox J, et al. Accurate proteome-wide label-free quantification by delayed normalization and maximal peptide ratio extraction, termed MaxLFQ. Mol Cell Proteomics 2014;13:2513–2526.

관련 글