LC-MS has become one of the most valuable analytical platforms in natural product research. It can reveal complex chemical profiles, support dereplication, and help researchers prioritize compounds before time-consuming isolation work begins.
This practical introduction explains the core LC-MS data-analysis workflow, the decisions that matter most, and the quality checks that make results easier to reproduce.
What LC-MS data actually contains
An LC-MS experiment produces more than a simple list of molecular masses. Each detectable feature is usually described by a retention time, a mass-to-charge ratio (m/z), and an intensity. Depending on the instrument and acquisition method, you may also have isotope patterns, adduct information, and MS/MS fragmentation spectra.
- Retention time helps distinguish compounds separated by chromatography.
- Accurate mass supports elemental-formula prediction and database searching.
- Peak intensity provides a semi-quantitative measure for comparing samples.
- MS/MS fragments provide structural clues and improve annotation confidence.
A practical LC-MS data-analysis workflow
1. Review instrument performance and quality controls
Before interpreting biological or chemical differences, confirm that the analytical system behaved consistently. Inspect blanks for carryover, pooled quality-control samples for signal stability, and internal standards for retention-time and intensity drift. A beautiful multivariate plot cannot rescue unreliable raw data.
2. Convert and organize raw data
Vendor files are often converted to open formats such as mzML before processing. Keep filenames systematic and preserve a sample sheet containing sample class, batch, extraction method, injection order, and dilution information. This metadata becomes essential when you troubleshoot batch effects or repeat an analysis months later.
3. Detect and align features
Feature-detection software identifies chromatographic peaks and groups related signals across samples. Important parameters include mass tolerance, minimum peak intensity, expected peak width, and retention-time alignment tolerance. Start with settings that reflect the actual instrument resolution and chromatographic performance rather than copying parameters from an unrelated study.
4. Remove artefacts and redundant signals
Raw feature tables commonly contain background ions, contaminants, isotopes, in-source fragments, and multiple adducts from the same compound. Blank subtraction, reproducibility filters, and adduct grouping can reduce this complexity. Document every filter so that another researcher can understand how the final table was produced.
5. Normalize and assess data quality
Normalization helps control for differences in sample amount, extraction recovery, or instrument response. The best method depends on the experiment and may use internal standards, sample mass, total useful signal, or quality-control-based correction. After normalization, examine feature distributions, missing values, QC variation, and principal-component analysis before testing biological hypotheses.
6. Annotate metabolites carefully
Accurate mass alone rarely proves a chemical identity. Strong annotations combine retention behavior, isotope evidence, adduct logic, MS/MS similarity, reference databases, literature evidence, and—when available—comparison with an authentic standard. Report annotation confidence transparently and distinguish confirmed compounds from tentative candidates.
Common mistakes to avoid
- Using processing parameters without checking whether they match the instrument and chromatography.
- Ignoring blanks, pooled QC samples, or injection order.
- Treating every feature as a unique metabolite.
- Reporting database matches as confirmed compound identities.
- Applying statistical tests before reviewing missing values, drift, and outliers.
Tools commonly used in natural product LC-MS research
Popular workflows may include software such as MZmine, MS-DIAL, XCMS, GNPS, SIRIUS, and vendor-specific platforms. Each tool has strengths and limitations. Choose software according to your data type, study goal, need for reproducibility, and access to computational support—not simply because a tool is fashionable.
Good LC-MS analysis is a chain of evidence. Reliable conclusions depend on transparent processing, quality-control evidence, and appropriately cautious annotation.
A simple starting checklist
- Define the biological or chemical question before processing.
- Review blanks, standards, pooled QCs, and injection order.
- Record every software version and processing parameter.
- Filter contaminants and unstable features before statistics.
- Use multiple lines of evidence for metabolite annotation.
- Save the raw data, metadata, scripts, and final feature table together.
With this foundation, LC-MS data becomes easier to interpret, compare, and reuse. Future Natural Product Hub guides will examine individual processing tools, molecular networking, annotation strategies, and reproducible reporting in more detail.