Skip to main content

Cause-and-Effect Relationships of Genes

A causal theory for studying the cause-and-effect relationships of genes


By studying changes in gene expression, researchers learn how cells function at a molecular level, which could help them understand the development of certain diseases.

But a human has about 20,000 genes that can affect each other in complex ways, so even knowing which groups of genes to target is an enormously complicated problem. Also, genes work together in modules that regulate each other.

MIT researchers have now developed theoretical foundations for methods that could identify the best way to aggregate genes into related groups so they can efficiently learn the underlying cause-and-effect relationships between many genes.

Importantly, this new method accomplishes this using only observational data. This means researchers don’t need to perform costly, and sometimes infeasible, interventional experiments to obtain the data needed to infer the underlying causal relationships.

In the long run, this technique could help scientists identify potential gene targets to induce certain behavior in a more accurate and efficient manner, potentially enabling them to develop precise treatments for patients.

“In genomics, it is very important to understand the mechanism underlying cell states. But cells have a multiscale structure, so the level of summarization is very important, too. If you figure out the right way to aggregate the observed data, the information you learn about the system should be more interpretable and useful,” says graduate student Jiaqi Zhang, an Eric and Wendy Schmidt Center Fellow and co-lead author of a paper on this technique.

Zhang is joined on the paper by co-lead author Ryan Welch, currently a master’s student in engineering; and senior author Caroline Uhler, a professor in the Department of Electrical Engineering and Computer Science (EECS) and the Institute for Data, Systems, and Society (IDSS) who is also director of the Eric and Wendy Schmidt Center at the Broad Institute of MIT and Harvard, and a researcher at MIT’s Laboratory for Information and Decision Systems (LIDS). The research will be presented at the Conference on Neural Information Processing Systems.

Learning from observational data

The problem the researchers set out to tackle involves learning programs of genes. These programs describe which genes function together to regulate other genes in a biological process, such as cell development or differentiation.

Since scientists can’t efficiently study how all 20,000 genes interact, they use a technique called causal disentanglement to learn how to combine related groups of genes into a representation that allows them to efficiently explore cause-and-effect relationships.

In previous work, the researchers demonstrated how this could be done effectively in the presence of interventional data, which are data obtained by perturbing variables in the network.

But it is often expensive to conduct interventional experiments, and there are some scenarios where such experiments are either unethical or the technology is not good enough for the intervention to succeed.

With only observational data, researchers can’t compare genes before and after an intervention to learn how groups of genes function together.

“Most research in causal disentanglement assumes access to interventions, so it was unclear how much information you can disentangle with just observational data,” Zhang says.

The MIT researchers developed a more general approach that uses a machine-learning algorithm to effectively identify and aggregate groups of observed variables, e.g., genes, using only observational data.

They can use this technique to identify causal modules and reconstruct an accurate underlying representation of the cause-and-effect mechanism. “While this research was motivated by the problem of elucidating cellular programs, we first had to develop novel causal theory to understand what could and could not be learned from observational data. With this theory in hand, in future work we can apply our understanding to genetic data and identify gene modules as well as their regulatory relationships,” Uhler says.

A layerwise representation

Using statistical techniques, the researchers can compute a mathematical function known as the variance for the Jacobian of each variable’s score. Causal variables that don’t affect any subsequent variables should have a variance of zero.

The researchers reconstruct the representation in a layer-by-layer structure, starting by removing the variables in the bottom layer that have a variance of zero. Then they work backward, layer-by-layer, removing the variables with zero variance to determine which variables, or groups of genes, are connected.

“Identifying the variances that are zero quickly becomes a combinatorial objective that is pretty hard to solve, so deriving an efficient algorithm that could solve it was a major challenge,” Zhang says.

In the end, their method outputs an abstracted representation of the observed data with layers of interconnected variables that accurately summarizes the underlying cause-and-effect structure.

Each variable represents an aggregated group of genes that function together, and the relationship between two variables represents how one group of genes regulates another. Their method effectively captures all the information used in determining each layer of variables.

After proving that their technique was theoretically sound, the researchers conducted simulations to show that the algorithm can efficiently disentangle meaningful causal representations using only observational data.

In the future, the researchers want to apply this technique in real-world genetics applications. They also want to explore how their method could provide additional insights in situations where some interventional data are available, or help scientists understand how to design effective genetic interventions. In the future, this method could help researchers more efficiently determine which genes function together in the same program, which could help identify drugs that could target those genes to treat certain diseases.

Genetic mutations, gene expression, molecular pathways, transcription factors, epigenetics, protein synthesis, signal transduction, gene regulation, genetic variants, gene-environment interaction, RNA splicing, chromosomal rearrangements, protein-coding genes, epistasis, gene editing, functional genomics, gene silencing, phenotypic traits, heritability, genetic disorders.

#Genetics #GeneExpression #MutationEffects #MolecularPathways #TranscriptionFactors #Epigenetics #ProteinSynthesis #GeneRegulation #GeneVariants #GeneEnvironment #RNA #ChromosomalChanges #ProteinCoding #Epistasis #GeneEditing #FunctionalGenomics #GeneSilencing #Traits #Heritability #GeneticDisorders

Comments

Popular posts from this blog

Genetics role in ovarian cancer

The Medical Minute: Genetics play big role in ovarian cancer In 2024, about 19,680 women in the United States will receive a new diagnosis of ovarian cancer and 12,740 women will die from the disease, said Dr. Shaina Bruce , a gynecologic oncologist at Penn State Cancer Institute . The median age of all patients who develop ovarian cancer is 63. Historically, women at increased risk for ovarian cancer are recommended to have their fallopian tubes and ovaries removed when they have completed having children. Taking that step to protect themselves comes at a heavy price ― surgical menopause. But Bruce said medical science is catching up with ovarian cancer. Studies could lead to new methods for preventative care and the surgery needed to lower risk may be easier than it once was. Below, during Gynecologic Cancer Awareness Month, Bruce discusses the disease and why acting to reduce your risk is worth it. What’s the connection between heredity and ovarian cancer? About 25% of all cases of ...

X chromosome

Gene on the X chromosome may help explain high multiple sclerosis rates in women Brain inflammation may be fueled by a gene on the X chromosome, a new study in mice suggests. And in female mice, who carry two X chromosomes, a diabetes drug called metformin may work to counteract that inflammation. If these findings bear out in later studies, they could help to unravel the long-standing mystery of why women, who have two copies of this inflammation-driving gene, are more prone to certain autoimmune diseases, particularly after menopause. A disparity between the sexes Our bodies are patrolled by immune cells that provide protection against bacteria and viruses, but sometimes, these defenses turn on us. In the autoimmune disorder multiple sclerosis (MS), for instance, the immune system attacks myelin, the fatty insulation surrounding the nerve fibers in the brain and spinal cord. This leads to symptoms such as muscle weakness and difficulty walking, as well issues with memory and thinking...

Multifactorial Genetic Conditions

Multifactorial Genetic Conditions Multifactorial genetic conditions are disorders caused by the combined effects of multiple genes and environmental factors , rather than a single gene mutation . These conditions do not follow classic Mendelian inheritance patterns and instead result from complex gene–environment interactions . Factors such as lifestyle, nutrition, infections, stress, and exposure to toxins can significantly influence disease onset and severity in genetically susceptible individuals. Common examples include diabetes, cardiovascular diseases , neural tube defects, asthma, and many neuropsychiatric disorders. Understanding multifactorial inheritance is essential for risk prediction, preventive medicine, and personalized healthcare strategies. Multifactorial inheritance, polygenic traits, gene–environment interaction, complex diseases, genetic susceptibility, environmental risk factors, non-Mendelian inheritance, disease predisposition, polygenic risk score, precision ...