Showing posts with label Molecular Biology. Show all posts
Showing posts with label Molecular Biology. Show all posts

Thursday, 18 June 2026

The Dark Proteome: 1,700 New Proteins Discovered in human DNA

The Dark Proteome: 1,700 New Proteins Discovered In Human DNA

The Dark Proteome: 1,700 New Proteins Discovered in Human DNA | Bioverse By Syed Muiz

Every few months, biology gets a new viral headline. This time the internet is flooded with claims like “Junk DNA has been proven wrong,” “Scientists discovered thousands of hidden genes,” “Biology textbooks must be rewritten,” and even “The biggest discovery since the Human Genome Project.” Many of these headlines were designed to maximise clicks rather than accurately explain the research. and AI-generated blogs start repeating simplified conclusions without properly explaining what the actual research paper really is. After reading the original Nature (2026) paper carefully, I realised that the study is important, but the real significance is different from the sensational version spreading online.

The researchers did not suddenly prove that all noncoding DNA is functional, nor did they rewrite molecular biology overnight. What they actually discovered is more technical and biologically interesting: thousands of previously overlooked genomic regions appear capable of producing small translated peptide products inside cells. Some may become recognised as functional microproteins in the future, while others may remain biologically uncertain.

Since most articles are oversimplifying the work, I decided to write this article directly from the research paper itself and explain what the study truly found, how scientists detected these hidden molecules, and why this research matters scientifically.

What Is the Dark Proteome?

The word proteome refers to—The complete collection of proteins produced by a cell, tissue, or organism. Traditionally, the human proteome was thought to consist mainly of proteins encoded by well-characterised genes.

But the dark proteome refers to—protein-like molecules that exist within cells but  difficult to detect, or absent from standard protein databases. These molecules often originate from unusual genomic regions that classical genome annotation methods overlooked. Simply, the dark proteome represents the unexplored part of the protein world inside living organisms.

Molecular biology focused mainly on well-characterised protein-coding genes. These genes contain sequences called ORFs (Open Reading Frames—these are some stretches of genetic sequence capable of producing amino acid chains). Our Classical genome annotation methods mainly recognised long and evolutionarily conserved ORFs because they resembled typical proteins.

But modern genomics reveals that many smaller and unconventional genomic regions also show signs of translation activity. These regions are called ncORFs (non-canonical Open Reading Frames — these are small overlooked genomic regions capable of translation but not traditionally recognised as normal protein-coding genes).

Many ncORFs exist inside non regions, inside untranslated RNA segments overlapping known genes, or within long noncoding RNAs. For years, these tiny sequences were often ignored because scientists assumed they were too short and biologically insignificant. But the new research challenges that assumption by showing that at least some of these regions generate detectable peptide products inside cells. 

What the New Research Actually Discovered

The researchers analysed 7,264 ncORFs using multiple experimental systems. Out of these, they found evidence for 1,785 translated peptide products through HLA immunopeptidomics. (HLA stands for Human Leukocyte Antigen )— molecules displayed on cell surfaces that present peptide fragments to immune cells. If peptides from ncORFs appear on HLA molecules, it means those regions were translated into amino acid inside cells.

"This is where scientific nuance becomes important. The paper clearly distinguishes between:  detecting the peptide products, and proving stable functional protein-coding genes."
These are not the same thing. Detecting translated peptides does not automatically mean scientists confirmed 1,785 brand-new fully functional proteins. Some products may indeed become recognised as genuine microproteins, while others may simply be unstable translation products, temporary stress-response peptides, or biologically unclear molecules. Because of this uncertainty, the researchers introduced the term: "Peptideins"

Peptideins

Peptidein is a new term introduced by the researchers for molecules that are clearly being produced inside cells through translation, but cannot yet be confidently classified as true proteins.  

In this study, scientists found thousands of previously overlooked genomic regions that are actively read by ribosomes and converted into small peptide products. The problem was that many of these molecules did not fit the traditional definition of a protein. Some may have important biological functions and could eventually be recognised as new proteins, while others may be short-lived molecules with limited cellular roles.
To avoid calling every translated product a protein without sufficient evidence, the Researchers and Author introduced the term peptidein. It acts as an intermediate category between noncoding sequences and fully established proteins.

The importance of this idea is that it challenges the old view that DNA regions are either coding or noncoding. Instead, the study suggests that there may be a spectrum of biological activity, with peptideins occupying the hidden space between these two extremes. This provides a new framework for exploring the dark proteome and understanding the thousands of newly discovered translated products reported in this research.

One of the strongest examples discussed in the paper was a peptidein called c10riboseqorf92 located inside a long noncoding RNA named OLMALINC. Long noncoding RNAs were traditionally believed not to produce proteins. Researchers used CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats—a genome editing technology used to selectively disable specific genetic regions) to disrupt c10riboseqorf92 and observe its biological effects. When this region was disrupted, survival decreased in many cancer cell lines, while important cellular pathways related to metabolism and DNA damage responses were also disturbed. The researchers observed significant effects in 415 out of 485 tested cancer cell lines, strongly suggesting that this peptidein may possess genuine biological importance.

However, even after observing these strong cellular effects, the researchers still avoided fully classifying c10riboseqorf92 as a conventional protein. The reason is that its precise role in normal healthy physiology remains uncertain. The authors remained scientifically cautious and continued referring to it as a peptidein rather than immediately calling it a fully established protein.

This careful distinction reflects one of the central themes of the entire study: modern biology is revealing that the boundary between coding and noncoding regions may be far more complex than previously assumed.

How Scientists Detected These Hidden Molecules

One reason this study became scientifically important is because researchers did not rely on a single experiment. They combined several advanced molecular techniques together.

The first major method was Ribo-seq (Ribosome Sequencing—a technique that identifies RNA regions actively being translated by ribosomes). Ribosomes are the molecular machines responsible for protein synthesis inside cells. If ribosomes repeatedly bind to a genomic region, it suggests translation may be occurring there.

The second method was MS (Mass Spectrometry—a technique used to identify peptide fragments by measuring molecular mass). This provided direct evidence that amino acid products physically existed inside cells. Detecting tiny proteins is technically difficult because small peptides degrade quickly and often escape conventional protein-detection systems.

The third major technique was HLA immunopeptidomics. This became one of the strongest sources of evidence in the study because it demonstrated that peptide fragments derived from ncORFs were being processed and presented on cell surfaces through HLA molecules.

The researchers also used CRISPR screening. By knocking out certain ncORFs, researchers observed how cells responded. Some knockouts affected cell survival, metabolism, DNA damage pathways, and cellular regulation, suggesting potential biological functions.

Another important system introduced in the paper was ORBL (Open Reading Frame Conservation by Length). Instead of only studying amino acid conservation, ORBL measures whether evolution preserved the reading frame structure itself — including start codons, stop codons, and frame integrity. This helped researchers identify ncORFs that may be evolutionarily preserved despite having rapidly changing amino acid sequences.

Together, these methods created a much stronger framework than previous studies because they combined—Translation evidence, peptide detection, immune presentation, evolutionary analysis, and functional screening.

Final Thoughts

After carefully analysing the original research paper, the most scientifically accurate conclusion is that biology may contain a much larger hidden layer of translated products than previously recognised. The study detected evidence for 1,785 ncORF-derived peptide products, but the researchers themselves remain cautious about classifying most of them as fully established proteins.

The greatest significance of this study is that it expands our understanding of how extensive hidden translation activity may be inside cells. For decades, molecular biology mainly focused on large conventional proteins. This research suggests that cells may also produce many smaller translated products that remained invisible because of technological limitations. Some of these peptideins may eventually become—cancer biomarkers, immunotherapy targets, regulators of cellular stress responses, or previously unknown microproteins involved in metabolism and gene regulation.

And perhaps the most fascinating part is that these hidden molecular signals were not newly created. They were already present inside cells all along. Science simply developed the tools capable of seeing them more clearly.

Wednesday, 6 May 2026

The Central Dogma Isn’t Broken

The Central Dogma Isn’t Broken

In this article, Syed Muiz explains the Central Dogma, summarises a recent research paper, and shows how the new findings do not actually break the Central Dogma.

A few weeks ago, a new research paper came out from the Department of Biochemistry, Stanford University, with the title — "Protein-templated synthesis of dinucleotide repeat DNA by an antiphage reverse transcriptase [Cite as: Deng et al., Science 10.1126/science.aed1656 (2026).]". And after that — within days, the YouTube algorithms caught fire. “Central Dogma BROKEN?”... bla bla bla!!! screamed thumbnails with red arrows and shocked faces. Educators, content creators, and even a few over-caffeinated teachers began circulating a distorted narrative—that bacteria had just invented a way to make DNA without a nucleic acid template, beating everything we thought we knew about biological information flow—the Central Dogma.

The reaction is weird and somewhat understandable. Because the actual discovery is genuinely strange, and even beautiful. But the claim that it breaks the Central Dogma is not just wrong—it misses the point entirely. I think this happens because people start creating their content just to beat the internet algorithm, but that is not considered ethical behaviour. Even I was shocked for a moment out of excitement—I thought that after a very long time in the field of biology, something ground-breaking had happened. But after checking the real research paper available on the internet, I found that even some famous educators and online coaching institutes copy-paste from popular magazines where even the writers are not familiar with science and research papers. That creates absurdity.

So instead of relying on secondary sources, I went directly to the original research paper and examined how scientists themselves are interpreting it. And my understanding is that, this research does not rewrite the foundation of the concept, but rather adds a new dimension to how we understand molecular templating. 

In this article, I will dissect what the Central Dogma actually says, what the new DRT3 research really found, what was the purpose of the research, what were the procedures and techniques used, what results they obtained, and finally, the answer to “Is the Dogma broken?”

What Is the Central Dogma?

To understand the whole situation, firstly we have to know that what is central dogma really is. Francis Crick, in year 1958 first laid out what he called the “Central Dogma” of molecular biology. He later refined it in 1970, and here is what he actually wrote: 

"The central dogma of molecular biology deals with the detailed residue-by-residue transfer of sequential information. It states that such information cannot be transferred back from protein to either protein or nucleic acid."

The central dogma was never a rigid, universal law of “DNA makes RNA, RNA makes protein”; instead, it was an informational restriction—a negative assertion—to explain the limitations of biological information transfer.

The claim that the Dogma is "broken" typically conflates and merges Crick’s principle with James Watson's later, simplified summary, "DNA → RNA → protein", which appeared in his 1965 textbook The Molecular Biology of the Gene.

The full picture is a set of three general transfers (DNA→DNA, DNA→RNA, RNA→protein) that occur in all cells, plus three special transfers (RNA→RNA, RNA→DNA, DNA→protein) that occur only under certain conditions—like reverse transcription in retroviruses. What the Dogma forbids is protein→nucleic acid transfer. That means once information has been translated into the amino acid sequence of a protein, that information cannot be used to specify the sequence of a nucleic acid. It is an empirical observation, as we know it.

Procedures & Techniques

In this research, the researchers used a combination of structural biology, biochemistry, genetics, and bioinformatics to study this system. First, they cloned two versions of the DRT3 system from Escherichia coli (a common laboratory bacterium). They expressed the full gene cluster using affinity tags (small protein tags added to help in purification), and then purified the complete ribonucleoprotein complex (a structure made of protein and RNA together).
And then, they solved its three-dimensional (3D) structure using cryo-electron microscopy (cryo-EM, a technique where samples are frozen and imaged to reconstruct high-resolution structures). They achieved a resolution of 2.6 Å. They captured two states of the system:

  • A resting state
  •  And an active state where DNA synthesis was happening in the presence of dNTPs (deoxynucleotide triphosphates, the building blocks of DNA)

To understand what kind of DNA this system produces, they performed in-vitro polymerase assays ( A lab experiments where enzyme activity is tested outside the cell) using different nucleotide combinations and mutated versions of the active site (the functional region of the enzyme).

After that, they analysed the DNA products using:

  • Next-generation sequencing (NGS, a high-throughput method to read DNA sequences) based on tagmentation (a process that fragments and tags DNA for sequencing)
  • And agarose gel electrophoresis


They also tested the biological function by performing phage infection assays (experiments where bacteria are exposed to viruses to check defence activity). In addition, they isolated escape mutants of phage T1 (virus variants that can bypass the defence system).

Finally, they carried out phylogenetic analysis (study of evolutionary relationships) across more than a thousand related bacterial systems to check how conserved these features are within the DRT3 family.

How They Got The Results

The results came from combining structural data with functional experiments. Cryo-EM showed that the whole system forms a D3-symmetric hexamer (a structure with six repeating units arranged symmetrically). It contains six copies of Drt3a, six copies of Drt3b, and six noncoding RNAs (ncRNAs).

In Drt3a, the mechanism was clear. The ncRNA has a conserved ACACAC sequence, which sits directly in the active site and acts as a template. From this, a growing DNA strand made of GT repeats (poly(GT)) was seen forming as an RNA–DNA hybrid (a duplex made of RNA and DNA).

And Drt3b was completely different. Its template-binding channel was physically blocked—so no RNA or DNA could enter. Still, it was producing DNA. Instead of using a nucleic acid template, the growing poly(AC) DNA strand was held in a bent and distorted shape inside the protein. This was controlled by multiple protein side chains (amino acid groups that interact with the DNA).

To confirm this, the researchers used mutagenesis (deliberately changing amino acids in the protein) along with in-vitro assays (lab-based enzyme activity tests). They found:

  • Drt3a alone produces only poly(GT) when given dGTP and dTTP (the DNA building blocks guanine and thymine), fully guided by the RNA template
  • Drt3b alone produces only poly(AC) when given dATP and dCTP (adenine and cytosine), even without any nucleic acid template


Then they mutated Glu26 (glutamic acid at position 26) to alanine or glutamine. This reduced accuracy and efficiency, and sometimes caused wrong nucleotides like dG (deoxyguanosine) to be added in place of dA (deoxyadenosine). This confirmed that specific amino acids in the protein act like a “template” by controlling which nucleotides are selected.

Finally, sequencing of the DNA product showed the expected pattern—alternating GT and AC strands forming a double helix. Phage experiments also revealed that a viral protein called ST61 (from phage T1) acts as a trigger, activating this defence system inside the bacterial cell.

What This Research Actually Explains

This study shows that a bacterial defence system can produce long, repetitive double-stranded DNA using two reverse transcriptases that follow completely different strategies. One strand is made in the usual way—by copying an RNA template.

And the other strand is made differently. It is guided by the protein itself. The enzyme’s active-site residues (specific amino acids inside the functional region) control which nucleotides are added, and forcing an alternating sequence—without reading any DNA or RNA template. This creates a new category of polymerase activity: DNA synthesis that is sequence-specific but template-independent (meaning it produces a defined pattern without copying an existing nucleic acid).

It also shows how evolution can modify ancient reverse transcriptase (RT, enzymes that convert RNA into DNA) structures to create new functions. Here, the protein structure itself enforces a strict dinucleotide pattern (a repeating unit of two nucleotides, like ACACAC) purely through its shape and chemical interactions.

At the biological level, this explains how the DRT3 system helps bacteria defend against phages (viruses that infect bacteria). And this study also identifies a viral protein called ST61 as the likely trigger that activates this defence system.

Central Dogma — Really Broken?

The answer is a clear No. Scientifically, what Drt3b is doing is genuinely new at the mechanistic level. No one has seen a reverse transcriptase use amino acid side chains to control nucleotide addition with this kind of alternating fidelity. That part is real and exciting but does it break the Central Dogma?

Here is the key point. The information that defines this alternating AC pattern is not coming from the protein in real time. It is already encoded in the gene that builds Drt3b. Glu26, Arg253 (arginine at position 253), Thr335 and Thr338 (threonine residues at positions 335 and 338)—all these critical amino acids are themselves products of a DNA sequence. They exist because the genome encoded them.

So what looks like protein → DNA is actually deeper than that. It is DNA → protein → DNA. A loop. Not a violation.

More importantly, the Central Dogma was never saying that proteins cannot interact with nucleic acids. That idea is incorrect. Proteins are constantly interacting with DNA and RNA—polymerases, helicases, transcription factors, all doing exactly that. The Dogma is not about interaction. It is about the flow of sequence information. And in that sense, Drt3b is not creating new information. It is executing a pre-set pattern.

Now the philosophical layer—what do we actually mean by “information”? If a protein’s fixed three-dimensional (3D) structure can guide nucleotide addition without a nucleic acid template, does that count as information storage or not?. But here is the important distinction—not every kind of guidance is information. In an RNA template, information exists because the sequence is read. There is a codon system, a mapping, a clear correspondence between sequence and output.

In the case of the protein, nothing like that is happening. The protein is not being “read” like an RNA template. There is no codon system here. No mapping where Glu26 corresponds to a nucleotide like dATP (deoxyadenosine triphosphate, a DNA building block). Instead, the protein creates a chemical environment inside its active site. This environment restricts what is possible. It allows only certain nucleotides to fit and react—mainly dA (deoxyadenine) and dC (deoxycytosine).

Thats mean it is simply creating a chemical environment, that allows certain nucleotides to fit and react, and rejects others.

So this is not coded, sequence-based information transfer. It is structural control or more directly—it is molecular forcing. That is why the term “protein-templated” is used very carefully. Normally, a template means complementarity—A pairs with T, C with G. But here, the protein is not acting as a complementary template. It is acting as a selector. And that selectivity is rigid.

And this philosophical shift is subtle. We now see that a protein’s three-dimensional structure can guide nucleotide addition over long stretches. This blurs the boundary between a catalyst and a template. But still—the Central Dogma isn’t broken.

More importantly—Drt3b cannot be reprogrammed to produce something else like poly(AG) instead of poly(AC). The moment key residues like Glu26 are changed, the system either loses stability or starts producing random sequences. So this is not a general system for encoding or transferring information. It just a highly specialised molecular machine.

The Real Scientific Position

So the Central Dogma stands — but it stands next to an open door. The door leads to a hallway of questions we had not thought to ask.

What other protein-templated syntheses exist in the vast and unexplored space of bacterial defense systems? And could a similar mechanism generate sequence-diverse DNA given a different set of templating residues? And if we can engineer such systems, what would that mean for our ability to write DNA without a template — truly de novo?

Those are the questions this paper leaves us with—not “Is the Dogma broken?” — that is a shallow and conceptually misleading headline. The real, unavoidable question is more precise and more unsettling: how many other biochemical mechanisms are we still blind to that can shape nucleotide sequences without actually transferring sequence-encoded information? Perhaps the mistake was never in the Dogma itself, but in treating it as a closed box rather than a living map.

If one day we discover a system that can read a protein’s amino acid sequence and accurately convert it back into DNA or RNA, then yes—that would break the dogma. This system does not do that. What it does instead is expand our understanding of what is possible inside the existing framework. It does not break the rule—it shows how much more exists within it. 

"This article is based on the paper “Protein-templated synthesis of dinucleotide repeat DNA by an antiphage reverse transcriptase” (Deng et al., Science, 2026, DOI: 10.1126/science.aed1656)."

Also published at: https://medium.com/@syedmuiz1

Friday, 2 January 2026

Rethinking Junk DNA: The Noise of the Genome

rethinking-junk-dna-noise-of-the-genome-human-plant-dna-experiment

Rethinking Junk DNA: The Noise of the Genome

A recent experiment in genomics did something that feels more philosophical than technical. Researchers took large stretches of plant DNA and placed them inside human cells. Not modified human sequences, not conserved regulatory regions, but foreign DNA from a plant that has no evolutionary relationship with humans. This DNA has never been selected, refined, or optimized to function inside a human nucleus. In theory, it should be meaningless in that environment. The simplicity of the setup hides a deep question: how much of what we observe in the human genome actually has biological meaning?. The logic behind the experiment was straightforward. For years, non-coding DNA in humans has been defended as functional because it shows biochemical activity. It is transcribed, bound by proteins, marked by chromatin modifications, and detected in multiple assays. But activity alone does not prove function unless we know what activity looks like in DNA that truly has no function. Plant DNA provides a natural control for this problem. Its sequences are effectively random from the perspective of human cells according to evolutionary biology it diverged over a billion years ago with no shared regulatory logic.

Researchers observed that, when the plant DNA was introduced into human cells, it did not remain silent. It became transcriptionally active. Human proteins bound to it. Chromatin marks appeared across its length. By many commonly used measurements, the plant DNA behaved very much like human non-coding DNA. Thats shows us that how noisy the system really is in broader perspective. well i discussed that topic in my article [The Dark Matter of Genetics: Junk DNA or Hidden Code?]

According to Biologists, Cells are not precise machines that only interact with meaningful sequences. They are crowded molecular environments where enzymes bind opportunistically and transcription machinery initiates wherever local conditions allow. RNA polymerase does not ask whether a sequence is evolutionarily important before engaging with it. If the physical properties are permissive, transcription can occur. This is not a failure of biology; it is a consequence of physics operating at the molecular scale. But activity and function are not the same thing. Likewise, A door swinging in the wind is active, but it is not opening for a reason. In the same way, DNA can be transcribed simply because the molecular environment allows it, not because the organism needs the product thats shows a deep and beautiful harmony of nature and life. The human–plant hybrid experiment makes this distinction very clearer.

This study does not argue that all non-coding DNA is useless. That would be just as incorrect as claiming that all of it is functional. Decades of genetics have clearly shown that some non-coding regions are essential. They regulate gene expression, organize chromosomes, guide development, and influence disease risk. Certain non-coding sequences are deeply conserved and thats shows organisms has more complexities then we assumed. What the experiment demonstrates is that background transcription is normal. When even foreign DNA shows similar levels of activity, it becomes clear that biochemical signals alone are a poor measure of biological importance. Function must be demonstrated through necessity, constraint, and consequence, not assumed from detection. This work also quietly corrects earlier overconfidence in genomics Because, At one point, widespread biochemical activity across the genome was interpreted as proof that most DNA is functional. Those claims were exciting but premature. Detecting activity is easy with modern tools. Proving that a sequence is required for survival, development, or reproduction is much harder. The human–plant hybrid cells redraw this boundary using experimental evidence rather than interpretation.

From an evolutionary perspective, the findings make sense. Evolution does not shows cleanliness or efficiency. It shows optimization for survival. If extra DNA does not impose a significant cost, there is little pressure to remove it. This explains why genome sizes vary so dramatically across species and why large amounts of repetitive or seemingly redundant DNA can persist for millions of years. Noise is tolerated because precision is expensive. The broader value of this research lies in clarity. It gives scientists a baseline for what non-functional DNA activity looks like inside living cells. It reminds researchers to be cautious when assigning meaning to signals. It encourages humility in a field that is often presented as complete when it is anything but. For the public, the message is balance. Dark DNA is not garbage. It is a mixed landscape shaped by evolution, physics, and time. Some regions matter deeply. Some matter indirectly. Many likely do not matter at all. Understanding which is which requires restraint, evidence, and patience.

Once this experiments comes into public domain, different people will cherry-pick it to suit their own agendas and try to prove their own narratives, But the genome is perfectly written script and It is a historical document filled with edits, leftovers, and annotations of varying importance. It is read by noisy molecular machines operating under physical constraints. Sometimes, placing a piece of plant DNA into a human cell reveals more about the nature of life than another layer of speculation ever could. But at this stage, drawing any conclusion would be premature. Because absence of evidence is not evidence of absence. After this experiment undergoes complete peer review- or rather, once it is academically complete- whatever results emerges will certainly broaden our perspective. But ii is unlikely that we will reach any final or definitive endpoint. Because with every veil that is lifted, it becomes even clearer how little we actually know when it comes to arrive at any ultimate conclusion.