Code
# Load required librariesAdvanced Modeling and Functional Analysis
In this problem set, you’ll work with the Brauer gene expression dataset to practice comprehensive tidyverse skills including data tidying, transformation, joins, pivoting, string manipulation, and statistical modeling using broom. The dataset contains gene expression measurements for yeast genes under different nutrient limitations and growth rates.
Before we start tidying and analyzing the data, take a moment to predict what you might find.
# Load required librariesTask 1: Load the raw Brauer gene expression data and examine its structure. What makes this data “messy” or untidy?
Breadcrumbs: Use read_tsv() to load the data from the URL. Examine column names and the first few rows. Think about tidy data principles - what issues do you see with the current format?
#|label: data-loading-02
# Load the Brauer gene expression data
url <- "https://github.com/rnabioco/molb-7950/raw/refs/heads/main/data/bootcamp/brauer_gene_exp_raw.tsv.gz"# Examine the structure of the dataWe want to create a tibble that looks like this:
# A tibble: 199,296 × 4
systematic_name nutrient rate exp_level
<chr> <fct> <dbl> <dbl>
1 YPR204W Glucose 0.05 1.17
2 YPR204W Glucose 0.1 1
3 YPR204W Glucose 0.15 0.86
4 YPR204W Glucose 0.2 0.77
5 YPR204W Glucose 0.25 0.53
6 YPR204W Glucose 0.3 0.3
7 YPR204W Ammonia 0.05 2.79
8 YPR204W Ammonia 0.1 2
9 YPR204W Ammonia 0.15 0.6
10 YPR204W Ammonia 0.2 0.16
# ℹ 199,286 more rows
# ℹ Use `print(n = ...)` to see more rowsNote the classes of each column - systematic_name is character, nutrient is a factor, rate is numeric, and exp_level is numeric.
Task 3: Transform the wide-format expression data into a long format suitable for analysis.
Breadcrumbs:
First select the relevant columns (systematic_name and the expression columns (G0.05, etc))
Then use pivot_longer() to convert expression columns to rows
The column names contain both nutrient type (G) and growth rate (0.05) information - use separate_wider_position() to split the first character (nutrient abbreviation) from the numeric rate. The key here is to specify widths = c(nutrient_abbr = 1, rate = 4) to split after the first character, and put the rest into rate. Not all the rates have 4 characters, so use too_few = "align_start" to handle that.
Create a nutrient lookup table to convert abbreviations to full names. The lookup table should look like this:
# A tibble: 6 × 2
nutrient_abbr nutrient
<chr> <chr>
1 G Glucose
2 N Ammonia
3 P Phosphate
4 S Sulfate
5 L Leucine
6 U UracilUse that table with a join function to add full nutrient names to the tidied data.
Convert the rate column to numeric and select the final columns in the desired order
Remove any NA values from systematic_name and exp_level using filter()
This is a “cheat” chunk to load the tidy data directly if you want to skip ahead. Otherwise, use the code above to create the tidy data from the raw data.
brauer_tidy_tbl <- read_tsv(
"https://github.com/rnabioco/molb-7950/raw/refs/heads/main/data/bootcamp/brauer_gene_exp_tidy.tsv.gz"
)
yeast_gene_info_tbl <- read_tsv(
"https://github.com/rnabioco/molb-7950/raw/refs/heads/main/data/bootcamp/yeast_go_terms.tsv.gz"
)
yeast_protein_props_tbl <- read_tsv(
"https://github.com/rnabioco/molb-7950/raw/refs/heads/main/data/bootcamp/yeast_prot_prop.tsv.gz"
)Task 6: Calculate summary statistics for gene expression by nutrient type.
Breadcrumbs: - Use group_by() and summarize() to calculate mean, median, and standard deviation of expression values for each nutrient. Which nutrients show the highest variability in expression?*
# Calculate summary statistics by nutrient# Find genes with extreme expression valuesInspect the results and note any patterns you observe in high and low expression genes across different nutrient-rate combinations.
Next, make a boxplot to visualize the distribution of expression levels for each nutrient condition. What insights can you draw from the plot?
#|Task 7: Find genes with extreme expression values under different conditions.
Breadcrumbs: For each nutrient-rate combination, identify the top 5 highest and lowest expressing genes. Use slice_max() and slice_min() or ranking functions. What patterns do you notice?
# Find genes with extreme expression values