Stops Mak­ing Sense

by Christoph

One of the most con­vinc­ing ar­gu­ments for a com­mon ori­gin of all known life forms on this planet is that they all use the same ge­netic code. In my un­der­stand­ing, the ar­gu­ment is strength­ened, per­haps some­what counter-in­tu­itively, by the nu­mer­ous find­ings that many life forms mod­ify the canon­i­cal ge­netic code in ways that suit them, and dis­play as­ton­ish­ing cre­ativ­ity in do­ing so. One re­cent ex­am­ple of how the ge­netic code can be tin­kered with made it onto the cover of the Jan­u­ary 26 is­sue of Na­ture (Fig­ure 1).
 

Fig­ure 1. Na­ture Cover from Jan. 23, Vol. 613, Is­sue 7945. Fron­ti­spiece: Blas­t­ocrithidia non­stop. F, fla­gel­lum; P, pro­mastig­ote stage; C, cyst-like straphanger stage (length ~3 µm). Photo Jan Votýpka. Source

In a screen for try­panoso­matid pro­tists − a fam­ily of plant, in­sect, and ver­te­brate par­a­sites of clin­i­cal and eco­nomic con­cern that in­cludes Try­panosoma cruzi, the causative agent of cha­gas dis­easeZáho­nová et al. (2023) iso­lated, cul­ti­vated, and se­quenced Blas­t­ocrithidia non­stop (see fron­tispiece). In the words of Kachale et al. (2023), au­thors of a fol­low-up study, Try­panosomatidae "are known for a wide range of odd­i­ties," and a par­tic­u­larly strik­ing "odd­ity" had caught Elio's at­ten­tion ear­lier, in Pic­tures Con­sid­ered #44 : The Ut­ter Magic of kDNA X. Brief­ly, kine­to­plast DNA (kDNA) of T. bru­cei is the mitochon­dri­al DNA of this pro­tist, which is ac­tu­ally a tan­gled DNA mesh­work of a few maxi‑circles and hun­dreds of mini‑circles. From a re­view by Jensen & Eglund (2005): "Maxi­cir­cle tran­scripts in  T. bru­cei are heav­ily edited, an amaz­ing re­action in which uridy­late residues are in­cor­po­rated into or re­moved from spe­cific in­ter­nal sites within the tran­script to cre­ate an open read­ing frame. Mini­cir­cles enco­de most of the small guide RNAs that are tem­plates for edit­ing speci­ficity." If you're think­ing of CRISPR now, you're not wrong.

Blas­t­ocrithidia non­stop had an­other "odd­ity" in store for the re­searchers. Záho­nová et al. (2023) found that the open read­ing frames of its genes are vir­tu­ally awash with in-frame stop codons (Fig­ure 2). Out of 7,259 pre­dicted pro­tein-cod­ing genes only 228 lack any in-frame UGA, UAA, or UAG stop codons. These are mainly genes that are highly ex­pressed and en­code cy­toso­lic ri­bo­so­mal pro­teins, trans­la­tion fac­tors, hi­s­tones, mi­to­chon­dr­ial elec­tron trans­port chain sub­units, and a few oth­ers. From the com­par­i­son with ho­mol­o­gous genes of other try­panoso­matids, the au­thors con­cluded that UGA has been re­as­signed to en­code tryp­to­phan (Trp, W), while UAG and UAA (UAR) en­code glu­ta­mate (Glu, E). Strik­ingly, UAA and, less fre­quently, UAG also seem to serve as bona fide stop codons. Be­cause pe­ri­ods within words are so un­usual, here is a "smell test":

I .ould b. sur­pris.d if you could r.ad this t.xt h.r. .ithout as­sis­tanc., but th. ri­bo­som.s of Blas­t­ocrithidia non­stop manag. to trans­lat. such "dott.d" mR­NAs .ffortl.ssly. Here's how they achieve this, in some more detail than explained in Figure 1.

UAA/UAG re‑assignment   The nu­clear genome of B. non­stop con­tains 70 tRNA genes, in­clud­ing tRNAGluCUA and tRNAGluUUA cog­nate to both UAR stop codons (R=G or A). A thor­ough phy­lo­ge­netic analy­sis by Kachale et al. (2023) sup­ports an ori­gin of these stop-codon-rec­og­niz­ing tR­NAs from stan­dard tR­NAsGlu, and ex­per­i­ments con­firmed the ex­pres­sion and charg­ing of both tRNAGluCUA and tRNAGluUUA in B. non­stop and their ab­sence in the try­panoso­matid model species T. bru­cei by North­ern blot analy­sis.

UAA is the pref­ered stop codon   Thus far, these au­thors have no real clue how trans­la­tional ter­mi­na­tion at UAA stop codons at the 3'‑ends of open read­ing frames (ORFs) can over­come readthrough by tRNAGluUUA. How­ever, and un­like in other try­panoso­matids, ORFs in B. non­stop ex­hibit a sig­nif­i­cant en­rich­ment of A in the cod­ing strand for the first ~40 bases down­stream of a stop codon (fol­lowed by a T‑stretch). This in­di­cates a pos­si­ble in­volve­ment of poly(A)-binding pro­tein (PABP), known to in­ter­act with A+U‑rich se­quences in other species.
 

Fig­ure 2. Scheme il­lus­trat­ing pos­si­ble con­se­quences of AT mu­ta­tional shift in Blas­t­ocrithidia an­ces­tor. Fre­quent GC-to-AT sub­sti­tu­tions in Blas­t­ocrithidia an­ces­tor could have led to the fol­low­ing con­se­quences that we ob­serve in B. non­stop genome: over­all AT-rich genome; ex­treme AT-rich­ness of in­ter­genic re­gions with fre­quent TAA codons; ap­pear­ance of in-frame stop codons. Pos­si­ble UAG-to-UAA and UGA-to-UAA sub­sti­tu­tions are evo­lu­tion­ary neu­tral, but al­lowed TGG-to-TGA, GAG-to-TAG, and GAA-to-TAA sub­sti­tu­tions. Source

UGA re‑assignment   A tRNA cog­nate to UGA en­cod­ing tryp­to­phan is miss­ing from the B. non­stop genome, and Kachale et al. (2023) found that the tRNATrpCCA does not un­dergo a CCA-to-UCA an­ti­codon edit­ing in the cy­tosol as it does in other try­panoso­matids. Sec­ondary struc­ture pre­dic­tions sug­gested that the ac­cep­tor stem (AS) of B. non­stop tRNATrpCCA is only 4 bp long, whereas the closely re­lated T. bru­cei and other trypanosoma­tids pos­sess a 5‑bp-long AS of the canon­i­cal tRNA length (see Fig­ure 3). A bat­tery of in vitro and in vivo tests in the cog­nate and in het­erol­o­gous sys­tems con­firmed that the "shorter" tRNATrpCCA does in fact lead to a sig­nif­i­cant in­crease in readthrough (=Trp in­cor­po­ra­tion) at in‑frame UGA codons.

The re‑assignment of three codons, in the case of the UGA re‑assignment in com­bi­na­tion with a mod­i­fied tRNA struc­ture, does not seem to be suf­fi­cient for Blas­t­ocrithidia non­stop to evolve and then seam­lessly cope with its "up­dated code."

First, Záhonová et al. (2023) had found in the pre­ced­ing study an un­usual Ser74Gly sub­stitution at a highly con­served po­si­tion in the trans­la­tional ter­mi­na­tion fac­tor eRF1. When Kachale et al.  (2023) in­tro­duced a sim­i­lar Ser67Ala sub­sti­tu­tion into the yeast or hu­man eRF1 ho­mologs and ex­pressed them in S. cere­visiae, they found a sig­nif­i­cantly in­creased readthrough at UGA but not at all at UAA/UAG codons. This demon­strates that Ser67Ala specif­i­cally re­stricts UGA de­cod­ing as stop codon in vivo, which is clearly a de­sir­able ef­fect in B. non­stop.

Sec­ond, in eu­kary­otes, the wide­spread non­sense-me­di­ated de­cay path­way (NMD) is re­spon­si­ble for the degra­da­tion of mR­NAs with pre­ma­ture stop codons as they are found through­out open read­ing frames in B. non­stop (Fig­ure 2). An aside: bac­te­ria have solved the re­lated is­sue of de­grading mR­NAs lack­ing a stop codon dif­fer­ently, by trans-trans­la­tion (see here in STC). One branch of the Trypanoso­ma­tidae has lost early in evo­lu­tion Upf1 and Upf2, the key com­ponents of the NMD path­way, and Kachale et al. (2023) as­sume that this might have been one of the pre­req­ui­sites of stop codon re‑assignment in B. non­stop.

Fig­ure 3. (A) Mod­i­fied nu­cle­o­sides, tRNA struc­ture and the strength of codon–anticodon in­ter­ac­tion in­flu­ence on codon de­cod­ing ac­cu­racy. tRNA de­cod­ing fi­delity is mainly mod­u­lated by the strength of the in­ter­ac­tion be­tween an­ti­codon and codon bases, tRNA abun­dance, mod­i­fied nu­cle­o­sides, and tRNA struc­ture. At the third codon po­si­tion (first an­ti­codon po­si­tion), there is a de­cod­ing flex­i­bil­ity that en­ables a sin­gle tRNA species to de­code more than one codon. This means that the 61 sense codons of the ge­netic code can be de­coded by <61 dif­fer­ent tRNA species. This de­cod­ing flex­i­bil­ity (wob­ble rule) is mod­u­lated by the na­ture of the first base of the an­ti­codon, par­tic­u­larly its mod­i­fi­ca­tion and also mod­i­fi­ca­tion of other bases in the an­ti­codon loop. Mod­i­fi­ca­tion of base 37 has a strong in­flu­ence on the fi­delity of de­cod­ing be­cause it mod­u­lates the in­ter­ac­tion of the third base of the an­ti­codon with the first codon base. The over­all struc­ture of the tRNA and, in par­tic­u­lar, the struc­ture of the an­ti­codon stem, also play an im­por­tant role in main­tain­ing de­cod­ing ac­cu­racy. Source. (B) tRNA-Phe from yeast show­ing mod­i­fied bases in blue m2G: 2‑methyl-guano­sine; D: 5,6‑Dihydrouridine; m22G: N2-di­methyl­guano­sine; Cm: O2'-methyl-cytdine; Gm: O2'-methyl-guanosine; T: 5‑Methyluridine (Ri­both­ymi­dine); Y: wyb­u­to­sine (Y‑base); Ψ: pseudouri­dine; m5C: 5‑methyl-cy­ti­dine; m7G: 7‑methyl-guano­sine; m1A: 1‑methyl-adeno­sine. CC BY-SA 3.0 Yikrazuul

A bit of his­tory  The long line of stud­ies of "al­ter­na­tive ge­netic codes" be­gan in the late 1970s when it be­came known that the ge­netic code of hu­man and yeast mi­to­chon­dria de­vi­ates from the canon­i­cal ge­netic code. San­tos et al. (2004) said in their re­view 20 years ago: "These stud­ies have re­vealed that the ge­netic code is still evolv­ing de­spite strong neg­a­tive forces work­ing against the fix­a­tion of mu­ta­tions that re­sult in codon re­as­sign­ment. Re­cent data from in vitro, in vivo and in sil­ico com­par­a­tive ge­nomics stud­ies are re­veal­ing sig­nif­i­cant, pre­vi­ously over­looked links be­tween mod­i­fied nu­cle­o­sides in tR­NAs, ge­netic code am­bi­gu­ity, genome base com­po­si­tion, codon us­age and codon re­as­sign­ment." And they pre­dicted that "...the study of ge­netic code vari­a­tion will prob­a­bly pro­vide im­por­tant new in­sights into how mRNA de­cod­ing fi­delity (con­trolled by the trans­la­tional ma­chin­ery) shapes genome base com­po­si­tion, codon us­age and re­as­sign­ment, and will ul­ti­mately re­veal novel mol­e­c­u­lar mech­a­nisms that link the en­vi­ron­ment to the evo­lu­tion of genomes." Vari­a­tions of the ge­netic code that in­volve re‑as­sign­ments of one or up to all three stop codons are known from di­nofla­gel­lates like Amoebo­phrya sp., and cil­i­ates like Ble­phar­isma sp., Par­duczia sp., and Condy­lostoma mag­num. The study by Kachale et al. (2023) has now added an­other piece to this great puz­zle (see Fig­ure 3 for de­tails).

Be­fore I drift into the meta­phys­i­cal, here is a prac­ti­cal ex­am­ple (a lit­tle ex­er­cise, if you like). Sup­pose you stum­ble upon the DNA se­quence of the rpoA gene en­cod­ing 'RNA Poly­merase sub­unit al­pha' of Ser­ra­tia sym­bi­ot­ica CWBI‑2.3, and you ask your­self how sim­i­lar the RpoA pro­teins of this Ser­ra­tia and the E. coli K‑12 ref­er­ence strain MG1655 may be. Be­cause, in gen­eral and with­out go­ing into the de­tails here, it makes more sense to cal­cu­late similarity/ho­mology on the pro­tein level than on the DNA se­quence level. You find the Ser­ra­tia RpoA pro­tein se­quence in the NCBI pro­tein data­base with the ac­ces­sion num­ber QLH63909.1. Be­fore you now BLAST the two pro­teins against each other − and you will come up with "98% iden­tity," no sur­prise as Ser­ra­tia and E. coli are first cousins in the or­der En­te­ro­bac­terales − you would read the "FEATURES" sec­tion in the header of the file QLH63909.1 and find therein im­por­tant in­for­ma­tion:

1. the line /coded_by="complement(CP050855.1:2924630..2925619)" tells you that the pro­tein se­quence was trans­lated from the DNA se­quence and not de­ter­mined by pro­tein se­quenc­ing. This is the rule in genome se­quenc­ing projects al­most with­out ex­cep­tion. A more re­cent ex­cep­tion I'm aware of is Su'etsugu et al. (2008), who de­ter­mined the cor­rect trans­la­tional ini­ti­a­tion site of the E. coli hda gene at the rare start codon CUG by pro­tein se­quenc­ing.

2. the line /transl_table=11 tells you that a spe­cific NCBI codon us­age ta­ble was ap­plied, click on the "11"; that's the one for En­ter­obac­terales. The same codon us­age ta­ble was used for the RpoA se­quence of E. coli, I checked that for you. To see, for ex­am­ple, the NCBI codon us­age ta­ble for Blas­t­ocrithidia, scroll down to "31". (or click "31" here, for con­ve­nience).

So, you can ac­tu­ally com­pare pears with ap­ples if you take the pro­tein se­quences with the cor­rect codon us­age ta­ble, but not the DNA se­quences. Learning/Exercise goals met 🗹
 

Stop Mak­ing Sense  is a mu­sic film from 1984 fea­tur­ing the US rock band Talk­ing Heads (see here the fa­mous re­lease poster). It had noth­ing to do with them at all, but re­searchers in­ves­ti­gat­ing trans­la­tional ter­mi­na­tion re­lated it to them­selves, and it first ap­peared as a ti­tle in PubMed in 1988:  Stop mak­ing sense: or Reg­u­la­tion at the level of ter­mi­na­tion in eu­kary­otic pro­tein syn­the­sis (Valle & Morch (1988)). Since then, this ti­tle came up in PubMed every other year, or so. It at­tracted Leoš Shiv­aya Valášek, co-cor­re­spond­ing au­thor of the Kachale et al. (2023) pa­per, and Kelly Krause, who jointly de­signed the cover for Na­ture shown in Fig­ure 1. Ap­par­ently, they could not do with­out it. I couldn't re­sist ei­ther, and have gone along with Leoš and Kelly in their choice of the plural of "stop" since all three stop codons are re‑assigned in the pro­tist Blastocrithi­dia non­stop. By the way, the species name is not prop­erly la­tinized ac­cord­ing to the rules, but is that not fit­ting?

 

Do you want to com­ment on this post? We would be happy about it! Please com­ment on mastodon or Bluesky.
 

Other Posts

  • A Snip­pet: How Do Long Bac­te­ria Snip Apart?

    by Elio — Surely there are lim­its to the size and phys­i­ol­ogy of bac­te­r­ial cells. They can only get to be so big or so small, so full of ri­bo­somes or so de­pleted of them, so fast or so ex­tremely slow in their growth. Add an­other one to this list: are there lim­its to how pre­cisely bac­te­ria po­si­tion their di­vi­sion site? What brings this up is that many of the bac­te­ria en­coun­tered in...

  • ...rosette is a rosette. (3|3)

    by Christoph — Among the Alpha­pro­teo­bac­teria, a num­ber of spe­cies from phy­lo­ge­neti­cal­ly dis­tant­ly re­lated or­ders are known for form­ing ro­set­tes dur­ing growth. De­spite com­mona­li­ties, they all have their own pe­cu­lia­ri­ties when it comes to the first cell‑cell con­tact(s) dur­ing "ro­set­ting". Phaeo­bac­ter in­hi­bens (or­der Rho­do­bac­ter­a­les) cells at­tach to each other by their "sticky ends" at one pole...

  • Bac­te­r­ial Blues

    by Janie — The color blue is an odd­ity. Of all pig­ments in na­ture – chem­i­cals that se­lec­tively ab­sorb and re­flect cer­tain wave­lengths of vis­i­ble light – blue ones are among the rarest. Pro­ducing blue dyes was once such a costly busi­ness that me­dieval Eu­ro­peans consid­ered it to be as pre­cious as gold...

  • Strep­to­myces spores tak­ing a ride...

    by Christoph — Strep­to­mycetes, these non-motile Acti­nobac­te­ria that are ca­pa­ble of mycelial growth much like fungi, are quite in­ven­tive when it comes to the dis­tri­b­u­tion of their spores. They don't just rely on en­chant­ing scents like geosmin that they emit to at­tract spring­tails, which trans­port their spores over long dis­tances (see Roberto's re­cent post). They also rou­tinely em­ploy bac­te­r­ial trans­port work­ers.

  • Mi­crobes and Methane

    by Roberto — If you're look­ing for a primer on mi­crobes and methane (as was I, re­cently) I di­rect you to a timely re­port from an Amer­i­can Acad­emy of Mi­cro­bi­ol­ogy (AAM) col­lo­quium on the sub­ject. It's ex­cel­lent read­ing as it con­tains a trea­sure trove of in­for­ma­tion along with an im­pres­sive list of ref­er­ences, for those who want to take a deeper dive into the sub­ject.