Proteomics

A MaxQuant Run Took Me 10+ Hours on a Windows PC — the Checks I Now Do Before Pressing Start

On an ordinary Windows PC, one MaxQuant run took over 10 hours. At that price every wrong setting costs a day. These are the mistakes I made — iBAQ/LFQ left off, FASTA, match-between-runs, unfiltered proteinGroups — and the checks that now come first.

·6 min read
#MaxQuant#MaxQuant mistakes#proteomics#LFQ#iBAQ#match between runs#proteinGroups.txt#Perseus#label-free quantification

MaxQuant pre-flight checklist

The Number That Changed How I Use MaxQuant

I ran MaxQuant on an ordinary Windows desktop, not a cluster. One run took more than 10 hours.

That number is the whole reason for this post. When a run takes ten hours, a wrong checkbox does not cost you a minute — it costs you a working day, and you usually find out the next morning. Most of what I got wrong was not difficult. It was small, it was silent, and it was expensive because of the clock.

The step-by-step MaxQuant tutorial on this site tells you how to set things up. This is the other half: what I set up wrong, how it showed up, and what I check now before pressing Start.

Mistake 1 — Leaving iBAQ or LFQ Off

The run finishes, proteinGroups.txt looks normal, and the quantification columns you needed are simply not there.

They are not on by default. LFQ is set per parameter group, under label-free quantification in the group-specific settings. iBAQ is a separate checkbox in the global settings. Both are easy to miss because nothing warns you — MaxQuant has no idea which numbers you intended to use downstream.

What I check now: before starting, I write down which column my downstream analysis will read (LFQ intensity … or iBAQ …), then go and find the setting that produces it. If I cannot point at the checkbox, I do not start the run.

Mistake 2 — The FASTA

Two separate ways to get this wrong.

Contaminants. MaxQuant can add its own contaminant list to the search. With it, keratins, trypsin and serum proteins come out labelled as contaminants and you can filter them. Without it, they come out as ordinary proteins and sit near the top of your abundance ranking, looking like findings.

Mixing species in one search. If your samples come from more than one organism and you merge the FASTA files into one search, tryptic peptides that are identical between species get assigned ambiguously, and the per-species quantification quietly breaks. Nothing errors; the numbers are just wrong. One search per species — the long version is in the shared-peptide problem.

What I check now: the FASTA list in the run, the organism of each file, and that the contaminant option is on.

Mistake 3 — Match-Between-Runs Without Thinking About It

Match-between-runs transfers identifications from one run to another using retention time and mass. It reduces missing values, which is why people turn it on. It also means some of your identifications were never fragmented in that sample — they were borrowed.

Whether that is acceptable depends on the design. Between replicates of the same sample type it is usually what you want. Between very different sample types, or different species, it can borrow identifications that are not really there.

The expensive part for me was not the setting itself. It was changing it after the fact, because changing it means another full run.

What I check now: I decide MBR on or off from the design before the first run, and write the reason in my notes, so I do not reopen the question at hour nine.

Mistake 4 — Analysing proteinGroups.txt Without Filtering It

proteinGroups.txt is not a clean result table. It contains three kinds of rows you almost always remove first:

  • Reverse — decoy hits used to estimate FDR. They are not proteins.
  • Potential contaminant — the contaminants from Mistake 2.
  • Only identified by site — proteins identified only through modified peptides.

Skip this and every downstream number — protein count, missing-value rate, the top of the volcano plot — is computed on a table that still contains rows that should not be there.

library(tidyverse)

pg <- read_tsv("combined/txt/proteinGroups.txt", show_col_types = FALSE)

clean <- pg %>%
  filter(is.na(Reverse) | Reverse != "+",
         is.na(`Potential contaminant`) | `Potential contaminant` != "+",
         is.na(`Only identified by site`) | `Only identified by site` != "+")

c(before = nrow(pg), after = nrow(clean))

Perseus does the same thing with its row filter on categorical columns. Either way, it is the first step, not an optional one.

Mistake 5 — Fighting the Windows Desktop Itself

Running on a normal Windows PC brings its own problems, separate from proteomics.

  • Paths. Long folder paths, spaces, and non-English characters in the raw-file or output path are a common source of failures. I now keep everything in a short path like D:\mq\project\.
  • Threads versus memory. More threads is not automatically faster. If the thread count is set beyond what the machine's memory can support, the run can fail late instead of early — which, at 10+ hours, is the worst possible time to fail.
  • The machine going to sleep. A desktop that sleeps or installs updates overnight will stop a run that was expected to finish by morning.

What I check now: short path, thread count matched to the machine rather than to the maximum, sleep and automatic restarts turned off for the night.

The Check Before the Check: Save the Parameters

The single habit that saved me the most time: save the parameters to mqpar.xml before every run.

It sounds trivial. It means that when a run is wrong, I can open the file and see exactly what was set, instead of trying to remember which checkbox I had clicked ten hours earlier. It also means the next run starts from a known state rather than from whatever the GUI happened to show.

The Two That Cost Me the Most

Looking back, two of these cost far more than the rest: Mistake 1 and Mistake 5.

Leaving the quantification setting off is the one you only discover after the run ends, so it always cost a full re-run. The Windows desktop problems were worse in a different way — they made runs fail partway through, so the hours already spent were lost and the whole run started again.

In the list below, the items marked ★ are the ones that cover those two mistakes. If you skip the rest, do not skip those.

My Pre-Flight List

What I go through now, before pressing Start:

  1. ★ Which quantification column will downstream read? Is the setting that produces it on (LFQ / iBAQ)?
  2. FASTA: right organism per file, one search per species, contaminants on.
  3. Match-between-runs: decided from the design, reason written down.
  4. ★ Short output path, no spaces or non-English characters.
  5. ★ Thread count sized to the machine's memory.
  6. ★ Sleep and automatic updates off.
  7. Parameters saved to mqpar.xml.

And after the run, before any statistics: filter Reverse, Potential contaminant and Only identified by site.

If the Run Time Itself Is the Problem

Ten hours per run is part of why I moved most new DIA work to DIA-NN. The MaxQuant vs DIA-NN comparison covers when each one makes sense, and the DIA-NN tutorial covers the setup. When I do run MaxQuant, I no longer start it without the list above.

관련 글