Proteomics

After DIA-NN 2.x: From report.parquet and pg_matrix to a Result You Can Publish (210 Public Runs, R)

I ran DIA-NN 2.3.2 on a public 210-run SWATH dataset and worked through the downstream half: which output file to start from, what the protein matrix already filters, what match-between-runs added, how much is missing, and why 210 runs were really n = 70.

·12 min read
#DIA-NN#DIA-NN 2.x#report.parquet#pg_matrix#downstream analysis#MaxLFQ#match between runs#missing values#technical replicates#proteomics R

DIA-NN downstream analysis

Where the First Tutorial Stopped

The DIA-NN complete tutorial takes you from raw files to a finished run. It ends where most tutorials end: DIA-NN prints its summary, writes its output files, and exits.

This post is the half after that — and instead of describing it in the abstract, I ran it on a public dataset so every number below can be checked.

The dataset: PRIDE PXD071192, human placenta, SWATH-MS on a TripleTOF 5600 (Ok et al., PLOS One 2025). 70 placentas, each measured 3 times — 210 runs. I am using it as a worked example of the downstream workflow, not re-testing the study's biological question.

The Run in Numbers

DIA-NN2.3.2 (Academia build)
Modelibrary-free: library predicted from FASTA (--gen-spec-lib --predictor)
FASTAUniProt human Swiss-Prot incl. isoforms + a cRAP contaminant list
Predicted library5,465,154 precursors · 20,533 proteins
Match-between-runson (--reanalyse)
Mass accuracyMS2 23 ppm · MS1 10 ppm (fixed)
FDR--qvalue 0.01
Machinelaptop, i9-14900HX, 16 threads, 63 GB RAM, Windows
Wall timeabout 930 minutes for all 210 runs (first pass done at 815 min)

If you are planning a cohort-sized DIA-NN run on a single machine, that last line is the one to budget for.

Start From the Right File — It Changed in 2.x

DIA-NN 2.x writes its main report as report.parquet, not report.tsv. Many tutorials and scripts online still read report.tsv; with a current version that file simply is not there.

The two outputs that matter downstream:

  • report.parquet — the long, precursor-level master table: one row per precursor per run, with the q-values and quantities. You need R's arrow (or Python's pyarrow) to read it.
  • report.pg_matrix.tsv — protein groups × runs, already wide, already filtered by DIA-NN — though not by the filter you might assume (next section). This is where most protein-level work starts.

Here is the real header of the 2.3.2 protein matrix from this run:

Protein.Group | Protein.Names | Genes | First.Protein.Description | N.Sequences | N.Proteotypic.Sequences | Rep1_TO_01.wiff | Rep1_TO_02.wiff | …

Six annotation columns, then one column per run (210 here). The run columns are actually full file paths — in this run C:/work/PXD071192/Raw_data\Rep1_TO_01.wiff, shortened above. That matters: code that picks sample columns with patterns like \.raw$ or ^/ finds zero columns on Sciex .wiff data or Windows paths. Select "everything that isn't an annotation column" instead. If a script you found online expects different annotation columns, it was written for another version — check your own header first:

library(tidyverse)

pg <- read_tsv("report.pg_matrix.tsv", show_col_types = FALSE)
names(pg)[1:6]
dim(pg)    # this run: 1,995 rows × 216 columns

What the protein matrix has — and hasn't — filtered

There is no q-value column in pg_matrix, so you cannot see from the matrix how it was filtered. The DIA-NN README I read describes the matrices as filtered at 1% FDR using global protein-group q-values. The 2.3.2 output from this run does not match that, and I only found out by checking it against report.parquet (2,200,577 rows × 71 columns, 230 MB):

Protein groups
Rows in pg_matrix1,995
… with Global.PG.Q.Value ≤ 0.011,759
… with Global.PG.Q.Value > 0.01236
Groups at global ≤ 0.01 in the report1,778 (exactly the number the log prints)

So about 12% of the matrix rows fail a global 1% protein-group FDR. The matrix row count tracks the run-specific protein-group q-value instead (1,998 groups at PG.Q.Value ≤ 0.01). If your methods say "proteins at 1% FDR", restrict the matrix to the global set yourself:

library(arrow)

rep <- read_parquet("report.parquet",
                    col_select = c(Run, Protein.Group, Q.Value,
                                   Global.PG.Q.Value, PG.MaxLFQ))

keep <- rep %>%
  filter(Q.Value <= 0.01, Global.PG.Q.Value <= 0.01) %>%
  distinct(Protein.Group)              # 1,778 in this run

pg_strict <- pg %>% semi_join(keep, by = "Protein.Group")

Reading only the columns you need (col_select) matters with a 230 MB file — all 71 columns load much more slowly.

What those 236 rows look like: they are sparse. The median one has a value in 17% of runs (versus 93% for the rows that pass), only 10 of them are present in at least half the runs, and together they hold 3.1% of the matrix's values. A valid-value filter removes most of them anyway — but they still inflate the protein count you report and the missing-value rate you describe.

🔴 Which q-value matters. For protein-level work the column to watch is the protein-group Global.PG.Q.Value, not only the precursor Q.Value. They are different levels of the same idea, and filtering on the wrong one changes your protein count. LLM-generated scripts suggest the precursor column here more often than you would expect — verify against the DIA-NN documentation.

PG.MaxLFQ — and where PG.Quantity went

For comparing protein levels across samples, use PG.MaxLFQ — the MaxLFQ algorithm builds protein ratios from peptide ratios across runs, which is what makes the numbers comparable between samples.

Older DIA-NN reports also had a PG.Quantity column, and older scripts use it. The 2.3.2 report has no PG.Quantity. Its protein-level quantities are PG.MaxLFQ, PG.TopN, and a PG.MaxLFQ.Quality score. A script that selects PG.Quantity will stop with a missing-column error.

The values in pg_matrix are PG.MaxLFQ — I checked 2,200 randomly chosen matrix cells against the report: zero mismatches.

The Number That Didn't Match

The run log says:

Protein groups with global q-value <= 0.01: 1778

The protein matrix has 1,995 rows. For a while that gap was the part of this post I could not explain; the table in the previous section is the answer. The log counts groups that pass the global protein-group filter (1,778, reproduced exactly from report.parquet). The matrix's row count sits close to the run-specific filter instead (1,995 vs 1,998) — and it includes 236 groups that fail globally while leaving out 19 that pass globally.

This is the most common question I get: "the protein count in the log doesn't match my table." It isn't supposed to. Report the number you actually analyzed, and say which filter produced it.

What Match-Between-Runs Actually Added

With --reanalyse, DIA-NN runs twice: a first pass with the predicted library, then a second pass using an empirical library built from the first. The log reports identifications per run for both passes:

Per run, median (210 runs)First passSecond pass (MBR)
Precursors at 1% FDR7,42610,514
Proteins (protein-level)9231,407

That is roughly +40% precursors and +50% proteins per run. In this run, the second pass took about 115 minutes on top of the first pass's 815.

The gain is real, and so is the caveat: identifications transferred between runs are inferred from other samples. With 210 runs of one tissue type that is usually what you want. With very different sample types in one run, be more careful about what MBR is filling in.

How Much Is Missing

Even after MBR, the matrix is far from complete — and how complete it looks depends on the FDR filter above:

pg_matrix as writtenGlobal PG q ≤ 0.01, contaminants removed
Protein groups1,9951,720
Missing cells26.5%19.3%
Value in all 210 runs473455
Value in ≥ 70% of runs1,2841,252

Proteins quantified per run (as written): min 876 · median 1,479 · max 1,846.

The 236 groups that fail the global filter account for most of the difference in missingness — another reason to apply it before you describe your data. The gap between ~1,700 groups and ~455 complete ones is the whole reason missing-value handling matters. The trade-offs between imputation methods are covered in missing values in proteomics and knn vs minDet vs MNAR.

The valid-value filter

Before any imputation or testing, decide which proteins are quantifiable at all:

logmat <- pg %>%
  select(Protein.Group, ends_with(".wiff")) %>%
  mutate(across(-Protein.Group, ~ log2(na_if(.x, 0))))

# one column per BIOLOGICAL sample (e.g. after averaging technical replicates — see below)
grp_a <- c("sample_A01", "sample_A02", "sample_A03")   # your column names
grp_b <- c("sample_B01", "sample_B02", "sample_B03")

valid_a <- rowSums(!is.na(logmat[grp_a])) >= 3
valid_b <- rowSums(!is.na(logmat[grp_b])) >= 3

quantifiable     <- logmat[valid_a & valid_b, ]
qualitative_only <- logmat[!(valid_a & valid_b), ]

A zero in a DIA-NN matrix means "not detected", not "zero abundance" — convert it to NA before log2, or every summary inherits -Inf.

⚠️ The pseudocount trap. If a protein is present in one group and absent in the other, the fold change is undefined. Adding a tiny pseudocount makes the arithmetic work and produces log2 fold changes of ±10 that dominate the volcano plot. I have watched a generated script do exactly this; the fix is the valid-value filter above, with one-sided proteins reported as presence/absence. The longer story is in DIA-NN to paper draft.

210 Runs Is Not n = 210

This dataset has three injections of each placenta. Those are technical replicates. The biological sample size is 70, not 210:

Group (from the sample annotation)PlacentasRuns
Depression with SRI treatment1854
Depression without SRI34102
Control1854

Treating the 210 runs as independent samples inflates n threefold and makes p-values look far smaller than they are. Two standard ways to handle it:

  • Collapse first — average the log intensities of each placenta's three runs, then test on 70 samples.
  • Model it — keep all runs and tell the model which runs come from the same placenta (in limma, duplicateCorrelation() with the placenta as the block).

Whichever you choose, write it in the methods. For how the choice of test behaves at different sample sizes, see limma vs DEqMS; for the full pipeline, Differential Expression Analysis in Proteomics.

Contaminants

This run searched a cRAP contaminant list alongside the human FASTA (with --cont-quant-exclude cRAP- set for the tagged entries). Even so, 48 protein groups in the final matrix carry the cRAP tag — DIA-NN does not drop them from the output for you. Remove them before analysis, and look at them once — they tell you about sample handling.

Whatever is most abundant in your prep can look like a finding. In a different reanalysis on this site, Matrigel nuclear proteins and smooth-muscle remnants topped the list for reasons that had nothing to do with the biology (PRIDE PXD023694 reanalysis).

Three Sanity Checks Before You Believe the Volcano

  1. Do the replicates cluster? Run PCA on the log2 matrix first. Here the three injections of each placenta should sit together; if they do not, fix that before testing anything.
  2. Are the top hits contaminants? Check your top 20 against the cRAP-tagged rows and the usual suspects for your prep.
  3. Does the direction survive a swap? Swap the two groups and rerun; every log2 fold change should flip sign and nothing else should move. It catches column-ordering mistakes before they become a reversed figure.

A Methods Paragraph You Can Adapt

Raw files were processed with DIA-NN 2.3.2 in library-free mode, using a spectral library predicted from the UniProt human Swiss-Prot database (including isoforms) supplemented with a cRAP contaminant list. Match-between-runs was enabled. Mass accuracy was fixed at 23 ppm (MS2) and 10 ppm (MS1). Precursors were filtered at 1% FDR and protein groups at 1% global protein-group FDR (Global.PG.Q.Value ≤ 0.01, applied to the protein matrix from the main report), and proteins were quantified with MaxLFQ. Contaminant proteins were excluded. Quantities were log2-transformed with zero values treated as missing. Technical replicates were [averaged per biological sample / modeled as a blocking factor], and proteins were retained when quantified in at least [k] of [n] samples per group.

The bracketed parts are the decisions that are yours to make for your own design.

FAQ

Should I start from report.parquet or report.pg_matrix.tsv?

pg_matrix for protein-level work — it is already wide. But use report.parquet alongside it to restrict the matrix to the global 1% protein-group FDR set, and whenever you need peptide-level evidence.

I can't find report.tsv.

DIA-NN 2.x writes report.parquet instead. Read it with arrow::read_parquet() in R or pandas.read_parquet() in Python.

Why does the protein count in the log differ from my matrix?

They use different filters. In this run the log's 1,778 is the global protein-group 1% FDR set; the matrix's 1,995 rows include 236 groups that fail that filter (and 48 contaminants). Filter the matrix to the global set from report.parquet, then report that number.

Do I need to re-run DIA-NN if I change a downstream threshold?

No — not for anything in this post. Re-running is needed only when you change what DIA-NN itself computed: the library, the FASTA, mass accuracy, or match-between-runs. At about 930 minutes for this dataset, that distinction matters.

Closing

Two things in this run surprised me, and both were invisible from the matrix alone: 12% of the matrix rows failing a global 1% protein FDR, and 210 runs that are really 70 samples. The first changes the protein count you report. The second changes every p-value.


Data: PRIDE PXD071192. Ok L. et al. (2025). Effect of depression and serotonin reuptake inhibitors antidepressant treatment during pregnancy on protein expression in the human placenta: A quantitative proteomics analysis. PLOS One 20(12): e0322090. doi:10.1371/journal.pone.0322090 (CC BY 4.0). Processing and all counts in this post are my own reanalysis of the public raw files.

관련 글