Explore the science of the human genome — from sequencing technologies to genetic disease, population genetics, and precision medicine.
Posted by GenomeReference · 38 replies
The human reference genome is a composite sequence assembled from multiple donors, serving as a standardized template against which individual genomes are compared to identify variants. The most widely used reference, GRCh38 (hg38), was released in 2013 and has been incrementally updated; the T2T-CHM13 assembly published in 2022 was the first truly gapless complete human genome. Every individual's genome differs from the reference at roughly 4 to 5 million positions, the majority of which are common variants with no known clinical significance. Rare, potentially pathogenic variants are typically identified by comparing an individual's variants to population databases and clinical variant repositories like ClinVar.
Posted by CopyNumberVariation · 43 replies
Copy number variations (CNVs) are structural genomic alterations in which a segment of DNA is duplicated or deleted relative to the reference, ranging from a few thousand to millions of base pairs. CNVs can affect gene dosage — the number of functional copies of a gene — and are associated with a range of developmental, neurological, and psychiatric conditions. For example, deletion of a segment on chromosome 22q11.2 causes DiGeorge syndrome, which affects heart development, immune function, and cognitive development. CNVs are routinely detected by chromosomal microarray analysis, which is recommended as a first-tier test for children with intellectual disability or developmental delay.
Posted by MicrobiomeGenetics · 35 replies
The gut microbiome — the community of trillions of bacteria, viruses, and fungi inhabiting the digestive tract — is influenced by both host genetics and environmental factors like diet and antibiotic use. Twin studies show that the heritability of specific microbial taxa ranges from 5% to over 40%, indicating a meaningful genetic component to microbiome composition. Host genetic variants affecting immune function, mucus production, and bile acid metabolism shape the intestinal environment in ways that favor or disfavor particular microbial species. Research is ongoing into how host-microbiome interactions affect predisposition to inflammatory bowel disease, metabolic syndrome, and mental health conditions.
Posted by StructuralVariants · 29 replies
Structural variants (SVs) include chromosomal rearrangements larger than 50 base pairs such as inversions, translocations, insertions, deletions, and copy number variations. They account for a large proportion of genetic diversity between individuals and contribute substantially to disease, yet are systematically underdetected by short-read sequencing technologies commonly used in clinical practice. Short reads of 150-300 base pairs struggle to span repetitive regions and complex rearrangements where SVs are concentrated. Long-read sequencing technologies from PacBio and Oxford Nanopore can span thousands to tens of thousands of base pairs, dramatically improving SV detection and enabling more complete genome characterization.
Posted by PopulationGenomics · 41 replies
Population genomics uses patterns of genetic variation across thousands of individuals from diverse populations to reconstruct ancient human migration routes, population splits, and admixture events. Studies of ancient DNA from archaeological specimens have revealed that modern humans interbred with Neanderthals and Denisovans after leaving Africa, with non-African populations carrying approximately 1-4% Neanderthal ancestry. The 1000 Genomes Project, gnomAD, and similar large-scale databases have catalogued common and rare genetic variation across major world populations, providing reference data for ancestry inference and medical genomics. Understanding population structure is critical for avoiding bias in genetic association studies and for accurately interpreting variant frequency data.
Posted by NonCodingDNA · 33 replies
Only about 1.5% of the human genome codes for proteins, but non-coding regions contain thousands of regulatory elements essential for controlling when, where, and at what level genes are expressed. Enhancers, silencers, promoters, and insulators are all non-coding elements that bind transcription factors and shape the gene regulatory landscape. Long non-coding RNAs (lncRNAs) regulate chromatin structure and gene expression through various mechanisms and are dysregulated in many cancers. The ENCODE project has systematically catalogued the functional elements of the human genome, demonstrating that at least 80% of the genome has some biochemical activity, though the significance of much of this activity remains actively debated.
Posted by AlzheimersRisk · 52 replies
The APOE gene has three common alleles: e2, e3, and e4. Carrying one copy of the APOE e4 allele roughly triples lifetime risk for late-onset Alzheimer's disease, while carrying two copies increases risk approximately 8-12 times compared to e3/e3 individuals. However, APOE e4 is neither necessary nor sufficient to cause Alzheimer's — many carriers never develop the disease, and many Alzheimer's patients do not carry the e4 allele. Rare early-onset familial Alzheimer's is caused by autosomal dominant mutations in APP, PSEN1, and PSEN2 genes, but these account for less than 1% of all cases. Genetic counseling is strongly recommended before and after testing for Alzheimer's-related variants due to the psychological complexity of results.
Posted by NewbornScreening · 36 replies
Newborn genomic sequencing pilot programs, currently being evaluated in several countries including the US and UK, aim to identify infants with treatable rare genetic conditions much earlier than traditional biochemical screening panels allow. Standard newborn screening covers 30-50 conditions depending on the state, while WGS has the potential to flag thousands of genetic conditions at birth. The BabySeq project and the BeginNGS initiative in the US are generating evidence on clinical utility, cost-effectiveness, and psychological impact of routine neonatal genome sequencing. Key challenges include managing incidental findings, parental consent complexity, and the vast majority of identified variants with unknown clinical significance.
Posted by TandemRepeats · 39 replies
Tandem repeat expansions are mutations in which a short DNA sequence repeated in tandem grows abnormally large across generations. Huntington's disease is caused by an expansion of a CAG trinucleotide repeat in the HTT gene beyond 35-40 copies, with longer expansions correlating with earlier disease onset — a phenomenon called anticipation. Myotonic dystrophy, fragile X syndrome, and several ataxias are also caused by repeat expansions in different genomic contexts and with different molecular mechanisms. Long-read sequencing technologies are particularly useful for characterizing repeat expansions because short reads often cannot span them. Researchers are actively developing therapies targeting pathological repeat RNA and the aberrant proteins they produce.
Posted by GenomicEthics · 44 replies
Sharing genomic data is essential for maximizing the scientific value of sequencing research, but it raises significant ethical concerns including privacy, consent, and potential misuse. Because a genome is uniquely identifying and contains information about relatives who may not have consented to sharing, traditional anonymization techniques are insufficient for genomic data. Controlled-access repositories like dbGaP and the European Genome-phenome Archive require data use agreements to limit access to qualified researchers. The Global Alliance for Genomics and Health (GA4GH) develops international frameworks for responsible genomic data sharing. Ongoing tensions between open science ideals and individual privacy rights are actively debated in bioethics and genomic medicine communities.
Join thousands of members sharing knowledge and experiences.