Pharmaceuticals · Europe

Making genomic analysis reproducible at research scale

A bioinformatics pipeline and platform programme to make analysis reproducible and independent of individual machines and scripts.

Genomic analysis pipelines running on scientific compute

Challenge

The challenge

Analysis pipelines had grown organically around individual researchers. Results depended on local environments, parameters were not captured consistently, and scaling sample volumes put growing pressure on shared infrastructure.

Context

The context

The scientific team had strong analysis expertise but limited software engineering capacity. The priority was to formalise what already worked, not to replace trusted tools.

Solution

Our approach

  1. 01 Captured existing pipelines and their tool and parameter dependencies
  2. 02 Reimplemented them as containerised, versioned workflows with recorded provenance
  3. 03 Introduced orchestration and scheduling across on-premise and cloud compute
  4. 04 Built a results interface for review and reporting rather than raw output directories
  5. 05 Established data and metadata management so runs remain comparable over time

Implementation

How it was delivered

Pipelines were migrated one at a time and reconciled against existing results, so scientific confidence in the new process was established empirically before old workflows were retired.

Let's engineer what's next.

Tell us about the problem. You will speak with an engineer who understands the domain, not a call centre.