Need Help?

Assembly of HiFi long-read sequencing for Korean whole-genome, K-001, K-002, K-003

Blood samples were collected from three Korean individuals (K-001: female, aged 50 years; K-002 and K-003: males, aged 29 and 36 years, respectively) at the Chungnam National University Hospital (Daejeon, South Korea), with written informed consent from all participants and approval from the Institutional Review Board (IRB number: CNUH 2019-06-034). All methods were performed in accordance with the relevant guidelines and regulations and carried out in accordance with the Declaration of Helsinki. Genomic DNA was isolated from 5 mL blood samples using DNeasy Blood & Tissue Kit (Qiagen, Carlsbad, CA, USA), according to the manufacturer’s instructions. The quality and quantity of the extracted genomic DNA were analyzed using an ND-1000 spectrophotometer (Thermo Fisher Scientific, Waltham, MA, USA). For long-read sequencing, we used the Sequel II HiFi system (Pacific Biosciences, Menlo Park, CA, USA). Briefly, HiFi sequencing libraries were prepared using a SMRTbell Express Template Prep Kit 2.0 (PN 101-853-100), followed by immediate treatment using a SMRTbell Enzyme Cleanup Kit (Pacific Biosciences). After pooling the fractions of the desired size (approximately 15–19 kb), the final libraries were further cleaned and concentrated using AMPure PB beads (Pacific Biosciences). Finally, library concentrations were assessed using a Qubit 1X dsDNA HS Assay Kit (Thermo Fisher Scientific), and the libraries were then sequenced using the Sequel IIe HiFi system (Pacific Biosciences). We assembled the raw HiFi reads de novo into contigs using hifiasm (version 0.16.0; default settings) and converted the GFA-formatted output to FASTA-formatted files.

Request Access

DUO:0000012
version: 2021-02-23

research specific restrictions

This data use modifier indicates that use is limited to studies of a certain research type.

DUO:0000021
version: 2021-02-23

ethics approval required

This data use modifier indicates that the requestor must provide documentation of local IRB/ERB approval.

DUO:0000042
version: 2021-02-23

general research use

This data use permission indicates that use is allowed for general research use for any research purpose.

Policy for the whole-genome data of Korea population

Whole-genome sequences of Korean populations generated by using long- and short-read sequencing technologies. We follow the relevant Korean laws and policies, such as "Act on the Acquisition, Management, and Utilization of Biological Research Resources" and "Bioethics and Safety Act".

Studies are experimental investigations of a particular phenomenon, e.g., case-control studies on a particular trait or cancer research projects reporting matching cancer normal genomes from patients.

Study ID Study Title Study Type
EGAS50000000375 Whole Genome Sequencing

This table displays only public information pertaining to the files in the dataset. If you wish to access this dataset, please submit a request. If you already have access to these data files, please consult the download documentation.

ID File Type Size Quality Report
Located in
EGAF50000315265 fa.gz 960.6 MB
EGAF50000315266 fa.gz 947.9 MB
EGAF50000315267 fa.gz 957.2 MB
EGAF50000315268 fa.gz 952.9 MB
EGAF50000315269 fa.gz 958.3 MB
EGAF50000315270 fa.gz 961.1 MB
6 Files (5.7 GB)