Skip to content

Learn derived attributes in peppy

This vignette will show you how and why to use the derived attributes functionality of the peppy package.

Derived attributes allow you to construct sample attributes dynamically using patterns and variables from other columns in your sample table. This is especially useful for building file paths without having to repeat directory structures for every sample.

Let's say you have samples with names like sample1, sample2, etc., and your data files follow a pattern like /data/{sample_name}.fastq. Instead of typing out the full path for each sample, you can use derived attributes.

Your sample table (sample_table.csv):

sample_name
sample1
sample2
sample3

Your project config (project_config.yaml):

pep_version: "2.0.0"
sample_table: sample_table.csv
sample_modifiers:
derive:
attributes: [file_path]
sources:
my_source: /data/{sample_name}.fastq

In your sample table, add the file_path column:

sample_name,file_path
sample1,my_source
sample2,my_source
sample3,my_source

When peppy loads this PEP, it will automatically expand my_source using the pattern from your config, replacing {sample_name} with the actual sample name. The result:

  • sample1 → /data/sample1.fastq
  • sample2 → /data/sample2.fastq
  • sample3 → /data/sample3.fastq

That's it! The my_source identifier gets replaced with the pattern defined in your config, and any variables in curly braces (like {sample_name}) get replaced with values from that sample's row.

The example below demonstrates how to use derived attributes in a more complex scenario with multiple data sources and multiple variables. Please consider the example below for reference:

examples_dir = "../tests/data/example_peps-cfg2/example_derive/"
sample_table_pre = examples_dir + "sample_table_pre.csv"
%cat $sample_table_pre | column -t -s, | cat
sample_name protocol organism time file_path
pig_0h RRBS pig 0 data/lab/project/pig_0h.fastq
pig_1h RRBS pig 1 data/lab/project/pig_1h.fastq
frog_0h RRBS frog 0 data/lab/project/frog_0h.fastq
frog_1h RRBS frog 1 data/lab/project/frog_1h.fastq

As the name suggests the attributes in the specified attributes (here: file_path) can be derived from other ones. The way how this process is carried out is indicated explicitly in the project_config.yaml file (presented below). The name of the column is determined in the sample_modifiers.derive.attributes key-value pair, whereas the pattern for the attributes construction - in the sample_modifiers.derive.sources one. Note that the second level key (here: source) has to exactly match the attributes in the file_path column of the modified sample_annotation.csv (presented below).

project_config = examples_dir + "project_config.yaml"
%cat $project_config | column -t -s, | cat
pep_version: "2.0.0"
sample_table: sample_table.csv
output_dir: "$HOME/hello_looper_results"
sample_modifiers:
derive:
attributes: [file_path]
sources:
source1: $HOME/data/lab/project/{organism}_{time}h.fastq
source2: /path/from/collaborator/weirdNamingScheme_{external_id}.fastq

Let's introduce a few modifications to the original sample_annotation.csv file to map the appropriate data sources from the project_config.yaml with attributes in the derived column - [file_path]:

examples_dir = "../tests/data/example_peps-cfg2/example_derive/"
sample_table = examples_dir + "sample_table.csv"
%cat $sample_table | column -t -s, | cat
sample_name protocol organism time file_path
pig_0h RRBS pig 0 source1
pig_1h RRBS pig 1 source1
frog_0h RRBS frog 0 source1
frog_1h RRBS frog 1 source1

Import peppy and read in the project metadata by specifying the path to the project_config.yaml:

from peppy import Project
p = Project(project_config)
p.sample_table
file_path organism protocol sample_name time
sample_name
pig_0h /Users/mstolarczyk/data/lab/project/pig_0h.fastq pig RRBS pig_0h 0
pig_1h /Users/mstolarczyk/data/lab/project/pig_1h.fastq pig RRBS pig_1h 1
frog_0h /Users/mstolarczyk/data/lab/project/frog_0h.fastq frog RRBS frog_0h 0
frog_1h /Users/mstolarczyk/data/lab/project/frog_1h.fastq frog RRBS frog_1h 1

As you can see, the resulting samples are annotated the same way as if they were read from the original, unwieldy, annotations file.