The dataset consists of - 126 whole exome sequencings (SAMD9/9Lmut: 64; GATA2mut 24, MDS wildtype 38/471) performed using SureSelect Human All Exon V6 enrichment (Agilent, cat# 5190-8863). The generated libraries were sequenced on the Illumina Hiseq 2500 with 150bp paired-end reads. FASTQ files were processed using SeqNext platform (JSI medical system, Germany), with gene-based alignment to a virtual panel of 300 genes (including 28 MDS-associated genes, SAMD9, and SAMD9L), consisting of genes relevant to bone marrow failure, MDS predisposition, and hematological cancers as per the Pan-Cancer studies with cohorts of >10,000 cancers. The respective BAM files are provided. - Custom panel targeting SAMD9, SAMD9L, and 22 single nucleotide polymorphisms (SNP) on chromosome 7q (allele frequency >35% in all ethnic sub-populations in gnomAD) (Ampliseq #IAD104171) were performed in 666/669 cases. And Custom panel targeting 28 MDS-associated genes (GATA2, RUNX1, HOXA9, CEBPA, GATA1, KRAS, NRAS, CBL, PTPN11, ASXL1, EZH2, SETBP1, FLT3, KIT, JAK2, JAK3, CSF3R, MPL, SH2, BCOR, BCORL1; RAD21, STAG2, CTCF, TP53, PTEN, CALR, VPS45) was performed in 544 cases (Ampliseq #IAD51150). Both custom panel libraries were prepared using NEBNext Ultra II DNA library prep kit (New England BioLabs, cat#E7645S/L) per manufacturer’s instruction and samples were sequenced on an Illumina Miseq 2000 with 2 x 150 bp reads. The respective BAM files are provided - 4 SAMD9/9L patients were subjected to MissionBio custom single-cell panel (CO-112) targeting 250 heterozygous gnomAD population polymorphisms on 7q arm and 69 amplicons in SAMD9/9L and other cancer genes. All libraries were sequenced on an Illumina NovaSeq6000 with 150 base-paired ending multiplexed runs. Fastq files were processed using the Tapestri Pipeline V2 and python-based Mosaic package (multi-omics analysis, data visualization). The derived BAM and loom files are provided.
This dataset contains 318 Tumor and Control WGS files submitted in another EGA box for samples for Gerhauser et al.,Cancer Cell, 2018, 34:996-1011. WGS and sequencing protocol was earlier described in Weischenfeldt et al, Cancer Cell, 2013.
To QC the TraCe-seq strategy, single-cell RNA-seq libraries were generated from a variety of human cancer cell lines transduced with the TraCe-seq library to validate the TraCe-seq strategy. Specifically, 5 different cell lines (PC9, MCF-10A, MDA-MB-231, NCI-H358, and NCI-H1373) were each transduced with a unique TraCe-seq barcode. The transduced cells were selected with puromycin only, dissociated to single cell suspensions, and then mixed together. The complex mixture of the 5 cell lines was profiled by 10X scRNA-seq. Furthermore, transduced NCI-H1373 cells were sorted by FACS to enrich for the top 50% of eGFP positive cells, and sorted cells were cultured briefly and used to construct scRNA-seq libraries and profiled by 10x scRNA-seq. To carry out the full TraCe-seq experiment, ~600 PC9 cells carrying unique TraCe-seq barcodes were expanded over 12 doublings to establish the barcoded population. A subset of the barcoded PC9 population was used to generate scRNA-seq libraries and profiled by 10x scRNA-seq prior to treatment to establish a baseline transcription profile for each barcoded clone. The rest of the cells were then treated for four days with 1 µM erlotinib, 1 µM GNE-069, or 1 µM GNE-104 respectively. scRNA-seq libraries were then generated form the treated cells and profiled by 10x scRNA-seq.
Dataset consists of Oncomine Comprehensive Cancer Panel v3 sequencing of 16 tumor-normal mucosa pairs. Tumors include 8 sessile serrated lesions (SSL), 3 sessile serrated lesions with dysplasia (SSL/D), 2 traditional serrated adenomas (TSA) and 3 tubular adenoma s(TA).
Pancreatic cancer biopsies and matching normal controls from 10 patients were exome sequenced. The same biopsies and PDX models derived from these were also subject to RNA sequencing.
This dataset contains 60 .bam files of shallow WGS data (~0.1X) from ovarian cancer cell lines. Sequencing reads were aligned to the 1000 Genomes Project GRCh37-derived reference genome using the BWA aligner (v.0.07.17; CRUK-CI alignment pipeline).
ATAC-seq profiling bam files from colorectal carcinoma and adenoma.
Transcriptomic data for five patients with breast cancer undergoing neoadjuvant chemotherapy and hyperpolarised 13C-MRI for early response assessment
SNP data for Ovarian cancer PRS (cases)
SNP data for 313 loci required for calculation of the Breast cancer PRS