Learn derived attributes in peppy
This vignette will show you how and why to use the derived attributes functionality of the peppy package.
-
basic information about the PEP concept on the project website.
-
broader theoretical description in the derived attributes documentation section.
What are derived attributes?
Section titled “What are derived attributes?”Derived attributes allow you to construct sample attributes dynamically using patterns and variables from other columns in your sample table. This is especially useful for building file paths without having to repeat directory structures for every sample.
Quick Start - Simple Example
Section titled “Quick Start - Simple Example”Let's say you have samples with names like sample1, sample2, etc., and your data files follow a pattern like /data/{sample_name}.fastq. Instead of typing out the full path for each sample, you can use derived attributes.
Your sample table (sample_table.csv):
sample_namesample1sample2sample3Your project config (project_config.yaml):
pep_version: "2.0.0"sample_table: sample_table.csvsample_modifiers: derive: attributes: [file_path] sources: my_source: /data/{sample_name}.fastqIn your sample table, add the file_path column:
sample_name,file_pathsample1,my_sourcesample2,my_sourcesample3,my_sourceWhen peppy loads this PEP, it will automatically expand my_source using the pattern from your config, replacing {sample_name} with the actual sample name. The result:
sample1→/data/sample1.fastqsample2→/data/sample2.fastqsample3→/data/sample3.fastq
That's it! The my_source identifier gets replaced with the pattern defined in your config, and any variables in curly braces (like {sample_name}) get replaced with values from that sample's row.
Advanced Example - Multiple Sources
Section titled “Advanced Example - Multiple Sources”The example below demonstrates how to use derived attributes in a more complex scenario with multiple data sources and multiple variables. Please consider the example below for reference:
examples_dir = "../tests/data/example_peps-cfg2/example_derive/"sample_table_pre = examples_dir + "sample_table_pre.csv"%cat $sample_table_pre | column -t -s, | catsample_name protocol organism time file_pathpig_0h RRBS pig 0 data/lab/project/pig_0h.fastqpig_1h RRBS pig 1 data/lab/project/pig_1h.fastqfrog_0h RRBS frog 0 data/lab/project/frog_0h.fastqfrog_1h RRBS frog 1 data/lab/project/frog_1h.fastqSolution
Section titled “Solution”As the name suggests the attributes in the specified attributes (here: file_path) can be derived from other ones. The way how this process is carried out is indicated explicitly in the project_config.yaml file (presented below). The name of the column is determined in the sample_modifiers.derive.attributes key-value pair, whereas the pattern for the attributes construction - in the sample_modifiers.derive.sources one. Note that the second level key (here: source) has to exactly match the attributes in the file_path column of the modified sample_annotation.csv (presented below).
project_config = examples_dir + "project_config.yaml"%cat $project_config | column -t -s, | catpep_version: "2.0.0"sample_table: sample_table.csvoutput_dir: "$HOME/hello_looper_results"sample_modifiers: derive: attributes: [file_path] sources: source1: $HOME/data/lab/project/{organism}_{time}h.fastq source2: /path/from/collaborator/weirdNamingScheme_{external_id}.fastqLet's introduce a few modifications to the original sample_annotation.csv file to map the appropriate data sources from the project_config.yaml with attributes in the derived column - [file_path]:
examples_dir = "../tests/data/example_peps-cfg2/example_derive/"sample_table = examples_dir + "sample_table.csv"%cat $sample_table | column -t -s, | catsample_name protocol organism time file_pathpig_0h RRBS pig 0 source1pig_1h RRBS pig 1 source1frog_0h RRBS frog 0 source1frog_1h RRBS frog 1 source1Import peppy and read in the project metadata by specifying the path to the project_config.yaml:
from peppy import Projectp = Project(project_config)p.sample_table| file_path | organism | protocol | sample_name | time | |
|---|---|---|---|---|---|
| sample_name | |||||
| pig_0h | /Users/mstolarczyk/data/lab/project/pig_0h.fastq | pig | RRBS | pig_0h | 0 |
| pig_1h | /Users/mstolarczyk/data/lab/project/pig_1h.fastq | pig | RRBS | pig_1h | 1 |
| frog_0h | /Users/mstolarczyk/data/lab/project/frog_0h.fastq | frog | RRBS | frog_0h | 0 |
| frog_1h | /Users/mstolarczyk/data/lab/project/frog_1h.fastq | frog | RRBS | frog_1h | 1 |
As you can see, the resulting samples are annotated the same way as if they were read from the original, unwieldy, annotations file.