MaxQuant 2.8 proteinGroups.txt, Scored Against a Known Answer: LFQ vs iBAQ vs Intensity, and What Match-Between-Runs Adds
I ran MaxQuant 2.8.1 on six public Orbitrap runs where human, yeast and E. coli were mixed at known ratios, then scored every intensity column against the right answer. LFQ was far more accurate than raw Intensity but quantified fewer proteins, iBAQ ratios were identical to Intensity, MBR made already-quantified proteins more accurate, and the old Reverse filter silently returns zero rows in 2.8.
Why Score MaxQuant Against a Known Answer
Most MaxQuant tutorials stop when proteinGroups.txt appears. The questions that matter come after: which of the three intensity columns to use, what the filter columns actually remove, and whether match-between-runs (MBR) adds real quantities or noise.
On real samples you cannot answer those, because nobody knows the true fold changes. So I used a public benchmark where the answer is known, ran MaxQuant 2.8.1 on it, and scored every column.
This is the MaxQuant counterpart to the DIA-NN 2.x downstream post. If you are setting up a run, the MaxQuant tutorial and the pre-flight checklist cover the part before this one.
The Dataset: Three Species at Known Ratios
PRIDE PXD028735 (Van Puyvelde et al., Scientific Data 2022) mixes human, yeast and E. coli digests in two fixed proportions:
| Human | Yeast | E. coli | |
|---|---|---|---|
| Condition A | 65% | 30% | 5% |
| Condition B | 65% | 15% | 20% |
| Expected log2(B/A) | 0 | −1 | +2 |
I used the Orbitrap Q Exactive HF-X DDA runs (120-minute gradient). The paper made each condition as three separate master batches. The file names carry Alpha, Beta and Gamma, which I take to be those batches, plus numbered repeat injections of each. I took one injection per batch (3 vs 3, 6 runs, 21.7 GB) so that repeat injections are not counted as replicates (the same trap as 210 runs that were really n = 70).
The Run
| Software | MaxQuant 2.8.1.0, MaxQuantCmd on Linux (WSL2 Ubuntu 24.04), .NET 8 runtime |
| FASTA | UniProt Swiss-Prot: human (20,431), yeast S288c (6,733), E. coli K-12 (4,531) + MaxQuant contaminants |
| Settings | Trypsin/P, 2 missed cleavages, Carbamidomethyl (C) fixed, Oxidation (M) + Acetyl (protein N-term) variable, 1% PSM and protein FDR, LFQ (min. ratio count 2), MBR on, iBAQ on |
| Machine | 16 of 32 threads, 62 GB RAM available to WSL |
| Wall time | 61.5 minutes for 6 runs · peak memory of the main process ~11 GB |
The slowest steps were retention-time alignment (9.2 min, part of MBR), the first search (6.7), mass recalibration (5.7) and the second-peptide search (5.6). Each run had 132,600–138,100 MS/MS spectra submitted, 29–30% of them identified, and 27,600–28,700 peptide sequences; 42,760 sequences across all six.
A first attempt with 28 threads crashed the whole WSL instance partway through; the most likely cause was the Windows drive that holds WSL's virtual disk running out of space. If you run MaxQuant under WSL, watch the free space on that Windows drive as well as what df reports inside Linux. The virtual disk also has its own size cap.
Gotcha 1: --create truncated my experiment names
MaxQuantCmd --create writes a parameter file from a raw folder. It named the experiment for LFQ_Orbitrap_DDA_Condition_A_Sample_Alpha_01.raw as A_Sample_Alph: prefix dropped, last letter cut. LFQ groups and labels everything by experiment name, so check the <experiments> block before running and set clean names yourself:
<experiments>
<string>A_Alpha</string>
<string>A_Beta</string>
...
</experiments>
Gotcha 2: There Is No Reverse Column Any More
In the 2.8.1 output, the decoy flag column is called Decoy. Scripts written for older versions usually filter on Reverse. In base R that filter does not fail. It silently empties your table:
pg <- read.delim("proteinGroups.txt", check.names = FALSE, quote = "")
nrow(pg) # 5785
nrow(pg[pg$Reverse != "+", ]) # 0 <- no error, no warning
subset(pg, Reverse != "+") # Error: object 'Reverse' not found
pg$Reverse is NULL, the comparison returns a zero-length vector, and indexing with it keeps nothing. If a downstream step then reports "0 proteins" you might look everywhere except the first filter. Check the column first:
decoy_col <- intersect(c("Decoy", "Reverse"), names(pg))
stopifnot(length(decoy_col) == 1)
What the Three Filters Remove
| Protein groups | |
|---|---|
Rows in proteinGroups.txt | 5,785 |
Decoy = "+" | 68 |
Potential contaminant = "+" | 59 |
Only identified by site = "+" | 44 |
| After all three | 5,628 |
| … human / yeast / E. coli / mixed-species | 3,780 / 1,457 / 389 / 2 |
E. coli is the smallest group, partly because its proteome is smaller and partly because it is only 5% of the material in condition A. Low-abundance proteins are the first to drop out, and that matters for every result below.
Gene names holds more than one name, separated by semicolons, in 172 groups (3.1%): the identified peptides cannot tell close homologs apart. Do not split those into separate rows (that double-counts the same intensity). Either keep the group label as it is, or report the first name and say so:
pg$gene <- sub(";.*", "", pg$`Gene names`) # first name only; say so in your methods
Which Intensity Column? The Scores
For each protein group quantified in at least two of three runs in both conditions, I computed log2(B/A) from the mean log2 intensity and compared it with the known answer. "Within ±0.5" is the share of proteins whose ratio landed within 0.5 log2 units of the truth.
| Column | Proteins scored (human / yeast / E. coli) | Within ±0.5: human (0) | yeast (−1) | E. coli (+2) |
|---|---|---|---|---|
Intensity | 3,603 / 1,250 / 243 | 86% | 74% | 43% |
iBAQ | 3,603 / 1,250 / 243 | 86% | 74% | 43% |
LFQ intensity | 2,896 / 907 / 132 | 99% | 92% | 81% |
LFQ is much more accurate, and quantifies fewer proteins. The spread (interquartile range) of the human ratios is 0.14 with LFQ versus 0.28 with raw Intensity; for E. coli it is 0.26 versus 0.95. But LFQ requires at least two peptide ratios per protein pair, so it leaves 20% fewer human and 46% fewer E. coli proteins quantified. Across all six runs, 74% of LFQ cells have a value versus 93.5% for Intensity. Which one you want depends on whether a missing protein or a wrong ratio hurts your study more.
iBAQ ratios are identical to Intensity ratios. iBAQ is the summed intensity divided by the number of theoretically observable peptides of that protein. The divisor is the same in every sample, so it cancels in any between-sample ratio. iBAQ is for comparing different proteins within one sample (rough absolute abundance); it adds nothing for group comparisons.
LFQ shifted the unchanged proteins. The human median under LFQ is +0.14, not 0. MaxLFQ normalisation assumes most proteins do not change between samples. This benchmark changes about a third of them on purpose (all yeast down, all E. coli up), and that is the likely reason the normalisation moved the unchanged ones. Raw Intensity, which is not normalised, put the human median at +0.04. In a real experiment with large, one-directional changes the same thing can happen without any warning, so check where your "unchanged" proteins sit before trusting small fold changes.
What Match-Between-Runs Added
MBR transfers identifications from runs where a peptide was sequenced to runs where it was only seen in MS1. peptides.txt records, per experiment, whether each value came By MS/MS or By matching.
Of 210,449 quantified peptide values, 20,830 (9.9%) came from matching. Among peptides quantified in all six runs, about 21% were complete only because of MBR. Scoring those peptides against the known ratios (centred on the human median):
| Within ±0.5 of truth | All 6 values by MS/MS | 1–2 runs by matching | 3+ runs by matching |
|---|---|---|---|
| Human (0) | 95% (n = 15,210) | 89% (3,078) | 91% (845) |
| Yeast (−1) | 93% (3,598) | 87% (844) | 91% (252) |
| E. coli (+2) | 88% (329) | 77% (116) | 62% (26) |
The matched values are centred on the right answer, so MBR is not inventing fold changes. They are noisier, though, and the noise grows where it hurts most: in low-abundance proteins that depend on several matched values. If a protein's significance rests on matched values in one condition only, look at it before you report it.
Protein level: the same run with MBR switched off
Peptide-level noise is only half the story, because MaxLFQ builds each protein ratio from many peptide ratios. So I re-ran the same six files with identical settings and MBR off. That took 52.0 minutes instead of 61.5; MBR cost about 9.5 minutes here, almost all of it retention-time alignment.
| LFQ intensity | Proteins scored, MBR on | MBR off | Within ±0.5, MBR on | MBR off |
|---|---|---|---|---|
| Human (0) | 2,896 | 2,346 | 99% | 99% |
| Yeast (−1) | 907 | 656 | 92% | 95% |
| E. coli (+2) | 132 | 77 | 81% | 62% |
LFQ cells with a value went from 60.5% (off) to 74.2% (on). MBR added 23% more scoreable human proteins, 38% more yeast and 71% more E. coli.
The two accuracy columns compare different protein sets, so the cleaner test is the proteins that were scoreable in both runs:
| Same proteins, LFQ | n | Within ±0.5, MBR on | MBR off | IQR on | IQR off |
|---|---|---|---|---|---|
| Human | 2,342 | 99.5% | 99.1% | 0.13 | 0.14 |
| Yeast | 656 | 97.7% | 95.0% | 0.15 | 0.18 |
| E. coli | 77 | 94.8% | 62.3% | 0.19 | 0.47 |
For the same low-abundance proteins, MBR made the LFQ ratios much better: with more peptides quantified in every run, MaxLFQ has more ratios to work with. The proteins that only became scoreable because of MBR are the weaker part. Of the 860 such protein groups, 96% of the human ones landed within ±0.5, but only 78% of the yeast and 62% of the E. coli ones.
For raw Intensity the effect was even larger, because a sum over fewer detected peptides is biased: within ±0.5 dropped from 86/74/43% (on) to 77/59/20% (off) for human/yeast/E. coli.
For this dataset, then: MBR improved the proteins you would have quantified anyway, and added new ones that are less reliable than the rest. If a result rests on a protein that is only quantified with MBR on, treat it as the weakest evidence in your table.
The Code
library(data.table)
pg <- fread("proteinGroups.txt", sep = "\t", quote = "")
decoy_col <- intersect(c("Decoy", "Reverse"), names(pg)); stopifnot(length(decoy_col) == 1)
pg <- pg[!(get(decoy_col) %in% "+") &
!(`Potential contaminant` %in% "+") &
!(`Only identified by site` %in% "+")]
A <- c("A_Alpha", "A_Beta", "A_Gamma"); B <- c("B_Alpha", "B_Beta", "B_Gamma")
lfq <- as.matrix(pg[, paste("LFQ intensity", c(A, B)), with = FALSE])
colnames(lfq) <- c(A, B) # columns are "LFQ intensity A_Alpha" etc.
lfq[lfq == 0] <- NA # MaxQuant writes 0 for "not quantified"
lfq <- log2(lfq)
ok <- rowSums(!is.na(lfq[, A])) >= 2 & rowSums(!is.na(lfq[, B])) >= 2
log2fc <- rowMeans(lfq[, B], na.rm = TRUE) - rowMeans(lfq[, A], na.rm = TRUE)
summary(log2fc[ok])
Two details that bite. MaxQuant writes 0, not NA, for a missing quantity; convert before log2, or you get -Inf. And use the same valid-value rule for every column you compare; otherwise you compare different protein sets.
A Methods Paragraph You Can Adapt
Raw files were processed with MaxQuant 2.8.1.0 against UniProt Swiss-Prot [organisms] and the MaxQuant contaminant list. Trypsin/P with up to two missed cleavages; carbamidomethylation of cysteine as fixed, oxidation of methionine and protein N-terminal acetylation as variable modifications; 1% FDR at PSM and protein level. Label-free quantification used MaxLFQ with a minimum ratio count of 2[, and match-between-runs was enabled]. Protein groups flagged as decoy, potential contaminant or only identified by site were removed. LFQ intensities were log2-transformed with zero values treated as missing, and proteins were retained when quantified in at least [k] of [n] samples per group.
FAQ
Should I use LFQ intensity or Intensity for differential expression?
In this benchmark LFQ was clearly more accurate, at the cost of fewer quantified proteins. For group comparisons, LFQ is the usual choice; keep raw Intensity in mind when low-abundance coverage matters, and check where unchanged proteins sit if you expect large, one-directional changes.
Is iBAQ better for fold changes?
No. Between samples it gives exactly the same ratios as Intensity. Use iBAQ to rank proteins by abundance within a sample.
Should I turn MBR on?
In this benchmark, yes: it made LFQ ratios of the same proteins more accurate (E. coli 95% vs 62% within ±0.5) and added 23–71% more scoreable proteins per species. The added proteins were less accurate than the rest, so treat MBR-only quantifications as the weakest evidence. This is a mixture of very similar samples; MBR is generally considered riskier when samples differ a lot.
My script filters Reverse and now nothing is left.
MaxQuant 2.8.1 calls that column Decoy. See Gotcha 2.
Data: PRIDE PXD028735. Van Puyvelde B, et al. A comprehensive LFQ benchmark dataset on modern day acquisition strategies in proteomics. Sci Data 2022;9:126. doi:10.1038/s41597-022-01216-6 (CC BY 4.0). Processing and all counts in this post are my own reanalysis of the public raw files.
Software: Cox J, Mann M. MaxQuant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification. Nat Biotechnol 2008;26:1367–1372. Cox J, et al. Accurate proteome-wide label-free quantification by delayed normalization and maximal peptide ratio extraction, termed MaxLFQ. Mol Cell Proteomics 2014;13:2513–2526.
관련 글
A MaxQuant Run Took Me 10+ Hours on a Windows PC — the Checks I Now Do Before Pressing Start
10월 1일 · 7 min read
ProteomicsFragPipe vs MaxQuant, Measured: Same Six Runs, Same Known Answer (9 Minutes vs 62)
10월 7일 · 9 min read
ProteomicsAfter DIA-NN 2.x: From report.parquet and pg_matrix to a Result You Can Publish (210 Public Runs, R)
10월 4일 · 12 min read
ProteomicsZero Proteins Passed FDR: How to Read and Report a Null Result (DIA-NN Reanalysis, 70 Placentas)
10월 4일 · 11 min read