On your marks, get set, go!

This is a lon­gread again (18 min). It dives deeper into "trans­la­tional ini­ti­a­tion" and is, there­fore, arranged like a mes­sen­ger RNA (mRNA) with proper ri­bo­some-bind­ing site (RBS), AUG start codon, open read­ing frame (orf), and am­ber (UAG) stop codon.

by Christoph

RBS
When did you first hear about the "ge­netic code"? In high school? Or in col­lege, in an undergradua­te course in mol­e­c­u­lar bi­ol­ogy and/or ge­net­ics? In ei­ther case you were cer­tainly shown the fa­mous codon ta­ble that spec­i­fies which of the 64 pos­si­ble triplets in mes­sen­ger RNA (mRNA) are transla­ted into one of the 20 canon­i­cal or "pro­teino­genic" L‑amino acids. I'm sure your teacher mention­ed that the ge­netic code is vir­tu­ally iden­ti­cal in E. coli  and the ele­phant, and, there­fore, prokary­o­tes and eukaryo­tes are much more closely re­lated than was imag­ined be­fore the code was eluci­da­­ted in the 1960s (~15 years be­fore Woese's "three do­mains"). Maybe you had to strug­gle, a lit­tle, to come to terms with de­gen­er­acy of the ge­netic code, mean­ing that it is re­dun­dant yet not ambi­guous: there are four codons for pro­line (P or Pro) but none of them codes for an­other amino acid. Maybe you were even told the fun story that the first iden­ti­fied "non-sense" or "trans­la­tional ter­mi­na­tion" or sim­ply stop codon (UAG) was called am­ber stop for Har­ris Bern­stein (Ger­man for am­ber), a grad stu­dent who helped with the ex­per­i­ments that led to its de­tec­tion. With a wink, the other two stop codons were called ochre (UAA) and opal (UGA). Lastly, your teacher prob­a­bly men­tioned that AUG codes for me­thio­n­ine (M, or Met) but is also used as "trans­la­tional ini­ti­a­tion co­don", in short: start codon. Which may sound triv­ial, but it's not.

Fig­ure 1. Di­a­gram for a set of 64 pET20b(+) plas­mids con­tain­ing medium-copy pBR322 ori­gin, IP­TG-in­ducible T7 pro­moter (T7) and superfol­der GFP re­porter (sfGFP). RBS, ri­bo­some-bind­ing site (5'-AGGAG); Spacer, op­ti­mal dis­tance (7 bp) be­tween RBS and the ini­ti­a­tion codon; T7 Term., T7 tran­scrip­tional ter­mi­na­tor; AmpR, am­­pi­cillin-re­sis­tance marker; pBR322, ColE1 plas­mid repli­ca­tion ori­gin. Source. Front­page: Im­age of an agar plate streaked with 16 diffe­rent strains of E.coli, each con­tain­ing a green flu­o­res­cent pro­tein with a dif­fer­ent start codon (an­no­tated along the edge of the plate). Com­pos­ite of two su­per-im­posed im­ages from a laser scan­ner. Credit: Jeff Glasgow/Ariel Hecht/Kelly Irvi­ne/ NIST. Source

AUG
Trans­fer RNAs (tRNA) are the ac­tual "si­mul­ta­ne­ous in­ter­preters" of the code pre­sented for trans­la­tion in a gene's tran­script, the mRNA. The tightly folded, L‑shaped tRNA mol­e­cules have two "busi­ness ends", an ex­posed 3‑letter "an­ti­codon" that matches the 3‑letter codon for an amino acid on one end, and the re­spec­tive amino acid cova­lent­ly linked to the other end. Yet to make sense of the nume­rous amino acid-loaded tRNA "words" by poly­mer­iz­ing them into a polypep­tide "sen­tence", a third type of RNA is re­quired, the ri­bo­some. Think of the ri­bo­some as a highly or­dered tan­gle of three RNAs adorned with a num­ber of pro­teins (55 in E. coli ), which forms pep­tide bonds bet­ween amino acids via its pep­tidyl­trans­ferase, a ri­bozyme arranged as a pocket within its 23S rRNA (in the 50S sub­unit, see a di­a­gram here). Ri­bo­somes slide along the mRNA that is "clamped" be­tween their two sub­units (30S + 50S in bac­te­ria) in a ratchet-like fash­ion. Specif­i­cally, they ex­pose a mRNA codon in their "A site" for bind­ing to the cog­nate tRNA while the grow­ing polypep­tide chain re­mains co­va­lently at­tached to the tRNA that had been pre­vi­ously bound to the A‑site and now, af­ter one move­ment of the gear, sits "next door" in their "P site" (see here for a schematic di­a­gram of the trans­lation cy­cle). Dur­ing on­go­ing trans­la­tion, ri­bo­somes move along mRNA in 5'→3' di­rec­tion in jolts of three bases and, since they don't hop back and forth, the trans­lated polypep­tide is co-lin­ear with the mRNA, from its N- to its C‑terminus (this is an over­sim­pli­fi­ca­tion, and there are well known ex­amples for pro­grammed frameshifts). When you add to this the fact that the ge­netic code lacks a "word di­vider" (a comma or a space) for the codon triplets, it is im­me­di­ately ob­vi­ous that ribo­so­mes bet­ter start trans­lat­ing mRNA in the "cor­rect" read­ing frame, be­cause the two other pos­si­ble read­ing frames can eas­ily re­sult in pep­tide gib­ber­ish. But how can ri­bo­somes dis­tin­guish be­tween AUG codons that sig­nal ad­di­tion of me­thio­n­ine (M, or Met) to a nascent pep­tide chain and those that are "meant" to be start codons?

Fig­ure 2. Trans­la­tion ini­ti­a­tion from all 64 codons. Nor­mal­ized per-cell fluore­scence mea­sured from three repli­cate cul­tures, grown in LB and re­sus­pended in PBS be­fore mea­sure­ment, with each of the 64 codons as the start codon in the GFP cod­ing se­quence. Shapes re­present the repli­cate plate num­ber, and filled shapes repre­sent GFP ex­pres­sion sig­nificantly greater (ad­justed P <0.05) than the non-ex­press­ing con­trol (Con­trol) as de­ter­mined by Dunnett's test. Note the log­a­rith­mic scale of the x‑axis. Source

E. coli solves this prob­lem by "di­ver­si­fi­ca­tion", and here's a brief ex­plainer. It's genome codes for two types of me­thio­n­ine-spe­cific tR­NAs, tRNAMet and tRNAfMet (mind the tiny "f", which means "formyl-"). Both have the CAU an­ti­codon, com­ple­men­tary to the AUG codon in mRNA, and both are "charged" with me­thio­n­ine by me­thionyl-tRNA syn­thetase, MetG. Af­ter charg­ing, Met-tRNAMet is "inten­ded for im­me­di­ate con­sump­tion" by trans­la­tional elon­gation, whereas Met-tRNAfMet is first con­verted by trans­formylase, Fmt, to fMet-tRNAfMet (N‑Formylmethionine-tRNAfMet). The N‑formyl group of the mod­i­fied methio­ni­ne pre­vents its use in pep­tide-bond for­ma­tion dur­ing trans­la­tional elon­ga­tion, but not its use for trans­la­tional ini­ti­a­tion. How­ever, an­other "di­ver­si­fi­ca­tion" is even mo­re im­por­tant. Al­though both me­thio­n­ine-spe­cific tR­NAs share the com­mon clover­leaf struc­ture of tR­NAs, their lack of sig­nif­i­cant se­quence ho­mol­ogy leads to the more-than-sub­tle dif­fer­ences in their re­spec­tive 3D struc­tures. As a con­se­quence of these struc­tural dif­fer­ences, Met-tRNAMet binds read­ily to elon­ga­tion fac­tor Tu (EF-Tu), and their joint com­plex fits neatly into the ri­bo­so­mal A site. In con­trast, fMet‑tRNAfMet has a high affin­ity for ini­ti­a­tion fac­tors 2 (IF‑2) and 3 (IF‑3), which co­or­di­nate the forma­tion of an "ini­ti­a­tion com­plex" com­pris­ing mRNA, a 30S ri­bo­so­mal sub­unit, ini­ti­a­tion fac­tors (IF‑1, IF‑2, and IF‑3), and fMet‑tRNAfMet (see here for a schematic di­a­gram sans the IFs). In this ini­ti­a­tion com­plex, the mRNA is "kept in place" through base-pair­ing with a short stretch of ho­mol­ogy to the 3' end of 16S rRNA (30S sub­unit), the ri­bo­some-bind­ing site (RBS, or "Shine-Dal­garno se­quence"), and ex­poses the AUG start codon ~7 nt down­stream on the mRNA (see here for a di­a­gram), the po­si­tion onto which IF‑1 and IF‑2 guide fMet-tRNAfMet. So, when the ini­ti­a­tion fac­tors dis­so­ci­ate from the initia­tion com­plex and the 50S sub­unit docks on, the ri­bo­so­mal P site is formed with the ini­tia­tor tRNA al­ready "in place" and the read­ing frame set.

Em­ploy­ing two dif­fer­ent AUG‑decoding tR­NAs for trans­la­tional elon­ga­tion and ini­ti­a­tion, respec­ti­vely, is cer­tainly an in­ge­nious so­lu­tion! (An aside: Ar­chaea and Eu­kary­otes were no less in­ven­tive and evolved dif­fer­ent mech­a­nisms.) How­ever, only 83% of all known E. coli genes have AUG start codons for the en­coded pro­tein. So, the nag­ging ques­tion re­mains why a dis­turbingly high per­cen­tage of genes have "al­ter­nate" start codons (14% GUG, 3% UUG, 1×AUU, 1×CUG). And E. coli is not par­tic­u­larly ec­cen­tric when it comes to start codons. In a num­ber of other bac­te­r­ial genomes the ra­tio of stan­dard vs. al­ter­nate start codons was found to be much the same. I say "dis­turbingly", be­cause ~15% al­ter­nate start codons can­not be sim­ply dis­missed as some weird "cod­ing er­ror" as was done when the GUG or UUG start codons for the N‑terminal me­thio­n­ine residues in the E. coli LacI (GUG), NusB (GUG), and Ndh (UUG) pro­teins were ex­per­i­men­tally con­firmed in 1974, 1981, and 1984, re­spec­tively (all 30+ years ago, think of that!).

orf
En­ter Ariel Hecht and col­leagues who asked in a re­cent study what would hap­pen to the transla­tion of a re­porter gene in which the start codon was re­placed by all 64 codons, one at a time? They con­structed multi-copy plas­mids with a green flu­o­res­cent pro­tein (GFP) as re­porter and whose tran­scrip­tion was dri­ven by a strong, in­ducible pro­moter. Trans­la­tion of the tran­scripts was trig­ger­ed by a RBS and, at the op­ti­mal dis­tance ('spacer'), a fully "per­mu­tated" set of 64 start codons (Fig­ure 1). These 64 plas­mids and a non-ex­press­ing con­trol plas­mid were in­tro­duced into E. coli (no, the plas­mids were not trans­formed into the cells, as the au­thors write, a mind-twist­ing habit found all too of­ten in pa­pers and heard in talks. Ac­tu­ally, phe­no­typ­i­cally an­tibi­otic-sen­si­tive cells were trans­formed by the added plas­mid DNA into  an an­tibi­otic-re­sis­tant phe­no­type that al­lowed se­lec­tion of plas­mid-car­ry­ing trans­for­mants (Av­ery et al. 1944)). Trans­for­mants were grown and in­duced un­der care­fully con­trolled con­di­tions in mi­crotiter plates. Fol­low­ing in­duc­tion, op­ti­cal den­sity and flu­o­res­cence were mea­sured in the cul­tures. In roughly two thirds of the cul­tures the flu­o­res­cence ex­ceeded auto-flu­o­res­cence of the non-ex­press­ing con­trol cul­ture (Fig­ure 2). The log­a­rith­mic scale used in this fig­ure for the x‑axis gives the vi­sual im­pres­sion of a smooth, al­most con­tin­u­ous vari­a­tion of the mea­sured flu­o­res­cence across all tested "start codons". In ac­tu­al­ity the val­ues sug­gest that the 64 codons would bet­ter be "binned" into four groups: three "canon­i­cal start codons" (AUG, GUG, UUG), from which trans­la­tion ini­ti­ated at 10 – 100% of AUG lev­els (and that to­gether ac­count for 99% of all start codons in E. coli, see above); four "near-cog­nates" (AUA, AUC, AUU and CUG), from which trans­la­tion ini­ti­ated at 0.1 – 1% rel­a­tive to AUG, 40 codons, from which trans­la­tion ini­ti­ated at 0.01 – 0.1% rel­a­tive to AUG, and 17 codons, from which trans­la­tion ini­ti­a­tion could not be de­tected at a level sig­nif­i­cantly above that of the non-ex­press­ing con­trol cells. Note that no trans­la­tion above back­ground oc­curred from the UAA and UAG (stop) codons, and only very weak trans­la­tion from the UGA (stop) codon, al­though in all three cases an interfe­rence with the pep­tide chain re­lease fac­tors RF1 (prfA ) and RF2 (prfB ) can be ex­cluded as these would rec­og­nize an "empty" ri­bo­so­mal A site and not an ini­ti­a­tion com­plex.

Fig­ure 3. Heat map for visua­lizing rel­a­tive trans­la­tion ini­ti­a­tion from each start codon, as mea­sured from ex­pres­sion of GFP in LB (see Fig­ure 2). The nor­mal­ized fluor­escence from three repli­cate cul­tures for each codon were av­er­aged, and plot­ted as the log10 of the av­er­age. Source

They took a dif­fer­ent look at the full range of their re­sults (from Fig­ure 2) by pro­ject­ing them onto a codon ta­ble (Fi­gure 3). Ap­par­ently, the strongest start codons have U as the sec­ond base. NAU (N=A,U,C,G) is an un­ex­pect­edly strong set of start codons, and G as the first base re­sults in stronger start codons than C as the first base in al­most all cases. They had checked be­fore (by pre­dic­tion) that the al­ter­nate start codons do not sig­nif­i­cantly tweak the mRNA's sec­ondary struc­ture, and that no other in‑frame start codons are present in the se­quence pre­ced­ing the ac­tual re­porter gene se­quence (an in‑frame GUG at the 16th codon in the GFP cod­ing se­quence would re­sult in a trun­cated, non‑fluorescent polypep­tide). They could there­fore safely con­clude that the trends they saw in the "al­ter­nate start codon ta­ble" re­flect all pos­si­ble vari­a­tions in codon::anticodon (CAU) ef­ficiency for trans­la­tional ini­ti­a­tion with­out be­ing blurred by too much noise of trans­la­tional er­rors. A good start­ing point for model build­ing in the Ångström scale and, much later, co‑crystalliza­tion ex­per­i­ments!

Fig­ure 4. Trans­la­tion ini­ti­a­tion from a sub­set of 12 co­dons span­ning the ex­pres­sion range. Trans­la­tion ini­tiated from three ex­pres­sion cas­settes, A GFP on a low-copy p15A plas­mid, B nanolu­ciferase on a low-copy p15A plas­mid and C nanolu­ciferase on a very-low-co­py BAC (miniF plas­mid). Tran­scrip­tion was dri­ven by the rhaP­BAD rham­nose-in­ducible na­tive E. coli pro­moter. Shapes rep­re­sent the repli­cate plate num­ber, and filled shapes rep­re­sent ex­pression sig­nif­i­cantly greater (ad­justed P < 0.05) than the non-ex­press­ing con­trol (Con­trol) as de­termined by Dunnett's test. Note the logarith­mic scale of the x‑axis. Source

Hecht et al. wor­ried that their re­sults with GFP expres­sion from a strong pro­moter on a high-copy plas­mid might have led to bi­ased re­sults by "over-stress­ing" the cell's trans­la­tional ca­pac­ity. In or­der to get closer to phy­siological con­di­tions they checked a set of 12 codons (from the "bins" men­tioned above) in ad­di­tional plas­mid con­structs, with an­other in­ducible pro­moter, an­other re­porter gene, and plas­mid back­bones with lower copy num­bers (Fig­ure 4). For var­i­ous ex­per­i­men­tal flaws that they don't fail to men­tion, none of these new con­structs worked sat­is­fac­to­rily (Fig­ure 4 A,B), ex­cept for the con­structs with a NanoLuc(iferase) re­porter gene on mini‑F plas­mids (1 – 2 copies per cell). For these con­structs, they mea­sured lu­mi­nes­cence val­ues that re­pro­duced the cor­responding val­ues found for the GFP re­porter on high-copy plas­mids (Fig­ure 2) al­beit at con­sid­er­ably lower sig­nal lev­els (Fig­ure 4 C). This re­sults en­sured that the first set of mea­sure­ments was in fact not bi­ased but I won­der why they did not choose chro­mo­so­mal in­te­gra­tion of their re­porter gene in the first place. This would have al­lowed them to test their con­structs un­der a full range of dif­fer­ent phys­i­o­log­i­cal con­di­tions with­out hav­ing to ac­count for vary­ing plas­mid co­py num­bers at dif­fer­ent growth tem­per­a­tures, for ex­am­ple (I'm aware that it's cheap to men­tion such tricks post fes­tum ).

Fi­nally, they set out to de­ter­mine via mass spec­trom­e­try how trans­la­tion of the re­porter gene be­gan at five se­lected codons with 1 – 3 "mis­matches" with re­spect to AUG (AUC, ACG, CAU, GGA and CGC (for the ex­perts: they used C‑terminal 6×His-tags for pro­tein pu­rifi­ca­tion and frag­mented the pro­teins us­ing Asp‑N en­do­pro­tease). They re­cov­ered sig­nif­i­cant amounts of pro­tein for all but the CGC vari­ant (the ex­pres­sion level was too low for pu­rifi­ca­tion), and all four had in­tact N‑termini, that is, ac­cord­ing to the pre­dicted se­quence, and, im­por­tantly, in­cluded an N‑terminal me­thio­n­ine. This con­firmed that fMet-tRNAfMet was in­deed used for ini­ti­a­tion, a re­sult in line with those obtai­n­ed for LacI and NusB ear­lier (see above). They al­most hid an im­por­tant re­sult in a half-sen­tence: "...In cul­tures with ACG as the start codon a small frac­tion of spec­tra (1 of 8) in­di­cated that the N‑terminal [amino acid of the] pep­tide might be the cog­nate amino acid, thre­o­nine (Mr = 119), with a mass shift of −30 Da rel­a­tive to me­thio­n­ine (Mr = 149)" (author's ad­di­tion in square brack­ets). This strongly sug­gests that free Thr-tRNAThr ("free"= not fully com­plexed by EF-Tu un­der con­ditions of gene ex­pres­sion from a multi-copy plas­mid) com­petes with fMet-tRNAfMet for bind­ing to the ACG start codon in the mRNA that is "kept in place" by the 30S ri­bo­so­mal sub­unit. Ap­par­ently, the per­fect codon-an­ti­codon in­ter­ac­tion for Thr-tRNAThr is strong enough to over­come the low affin­ity of the ini­ti­a­tion fac­tors IF‑1 and IF‑2 for non-ini­tia­tor tR­NAs with a prob­a­bil­ity of ~1:8. Con­versely, this means that dur­ing ini­ti­a­tion com­plex for­ma­tion, the strong in­ter­ac­tion of ini­ti­a­tion fac­tors IF‑1 and IF‑2 with fMet-tRNAfMet is suf­fi­cient to over­come the im­per­fect codon-an­ti­codon in­ter­ac­tion for the AUC, CAU, and GGA codons, and, at a ra­tio of 8:1, also for ACG (but not for CGC). So, it might ac­tu­ally be bet­ter to un­der­stand the "ge­netic code" not as a ma­trix of bi­nary yes/no de­ci­sions but each codon::anticodon:amino acid com­bi­na­tion as an in­te­gral of the affini­ties of all mol­e­cules in­volved. And "all mol­e­cules" does not only in­clude mRNA, tRNA and ri­bo­some, but elon­ga­tion and ini­ti­a­tion fac­tors as well. "All mol­e­cules" are highly dy­namic dur­ing trans­la­tion, and the con­for­ma­tional changes (me­chan­i­cal work) are dri­ven by GTP hy­drol­y­sis, which of course re­sults in vary­ing affini­ties through­out the en­tire process. This sounds com­pli­cated but look­ing at the cher­ished "codon ta­ble" from an­other per­spec­tive seems timely.

Hecht et al.'s ex­per­i­ments may seem like mere dili­gence, but, as they point out, their re­sults have more un­ex­pected con­se­quences than just ques­tion­ing the way we un­der­stand the ge­netic code: "Av­er­age per-cell abun­dances of pro­teins in bac­te­ria and mam­malian cells span five to seven or­ders of mag­ni­tude. Given that the non-canon­i­cal trans­la­tion ini­ti­a­tion shown in this pa­per spans about four or­ders of mag­ni­tude, it is pos­si­ble that this [vari­ance in] level of ex­pres­sion could be phys­i­o­log­i­cally sig­nif­i­cant and may serve as an ad­di­tional mech­a­nism for con­trol­ling pro­tein syn­thesis" (author's ad­di­tion in square brack­ets). In­deed, this doesn't seem far-fetched. Reg­u­lat­ing pro­tein ex­pres­sion by tweak­ing the mRNA's start codon would add to the no­tion that cells fol­low a vari­a­tion of Murphy's law when it comes to con­trol­ling pro­tein syn­the­sis: every­thing that can be reg­u­lated will be reg­u­lated. His­tor­i­cally, the vari­a­tion of pro­moter strength by se­quence altera­tions and reg­u­la­tion of pro­moter ac­tiv­ity by tran­scrip­tion fac­tors was ob­served first, think lac ope­ron. Since then it has be­come clear that, in ad­di­tion, mRNA half-life (codon-choice de­pen­dent mRNA sec­ondary struc­ture de­ter­mines sRNA and RNase ac­ces­si­bil­ity), mRNA "trans­lata­bil­ity" (codon-choice de­pen­dent trans­la­tion in­flu­ences the "drain" rate of tRNA pools), and pro­tein half-life (amino acid-se­quence de­pen­dent 3D fold­ing de­ter­mines pro­tease ac­ces­si­bil­ity) con­trol the ex­pres­sion level of a pro­tein – and, of course, this reg­u­la­tory net "folds & stretches" with the cell's phys­i­ol­ogy (tem­per­a­ture, growth rate). That growth tem­per­a­ture plays an im­por­tant role in gene ex­pres­sion is old hat, but in­di­rect ev­i­dence sug­gests that this also ap­plies to trans­la­tional ini­tia­tion: a mu­tant rIIB gene of phage T4 has the wild-type AUG start codon re­placed by AUA, which re­sults in re­duced trans­la­tion at 37°C in vivo and in vitro (to 10 – 15% of the wild-type level) and abol­ishes it com­pletely at 42°C.

And there are even evo­lu­tion­ary im­pli­ca­tions, again in their own words: "...there may be evolu­tio­nary util­ity to trans­la­tion ini­ti­a­tion from non-canon­i­cal start codons. Re­search with yeast has shown grad­ual tran­si­tions of ge­netic se­quences be­tween genes and non-genic ORFs in re­lated species. We can imag­ine a sce­nario wherein, over evo­lu­tion­ary time scales, point mu­ta­tions could cre­ate a weak non-canon­i­cal ini­ti­a­tion codon down­stream of a RBS. The small amounts of pro­tein pro­duced from such an ORF, if ben­e­fi­cial to the or­gan­ism, could se­lect for fur­ther mu­ta­tions that in­creased trans­la­tion ef­fi­ciency up to a point where the gene prod­uct more di­rectly im­pacted or­ganismal fit­ness." (see here in STC for an ex­am­ple of step-wise de novo  gene evo­lu­tion in S. ce­re­visiae ).

Fig­ure 5. Align­ment of the N‑terminal 51 amino acid residues for 26 ran­domly cho­sen DnaA pro­tein se­quences of the En­ter­obac­te­ri­aceae. The ClustalW align­ment was made by the au­thor us­ing the CLC se­quence viewer soft­ware pack­age.

UAG
Find­ing the "right" ini­ti­a­tion codon in an mRNA is also the task of bioin­for­mati­cians who comb through piles of DNA se­quence files when an­no­tat­ing pro­tein-cod­ing genes in newly se­quenced genomes. They have long since given up do­ing this by eye and use "an­no­ta­tion pipelines" in­stead, that is, soft­ware pack­ages that find open read­ing frames (orfs) in se­quence con­tigs, among other things. State‑of‑the‑art pipelines can deal with known vari­ants of the ge­netic code, for ex­am­ple, the code for ver­te­brate mi­tochondria. How­ever, even ad­vanced pro­grams mostly fall short of de­tect­ing "ex­otic" start codons like AUU – the E. coli gene for IF‑3, infC, con­tains an AUU start codon – but they rou­tinely find the "stan­dard" rare ini­ti­a­tion codons (GUG, UUG) in bac­te­r­ial genomes. Not al­ways, though, and here's an ex­am­ple. When you align en­ter­obac­te­r­ial se­quences of the (highly con­served) repli­ca­tion ini­tia­tor pro­tein, DnaA, you will see vir­tu­ally per­fect se­quence con­ser­va­tion but flut­ter­ing N‑termini in two cases (Fig­ure 5). For the En­ter­obac­ter can­cero­genus dnaA gene, GTG is an­no­tated as the ini­ti­a­tion codon (GUG=V, or Val, see the codon ta­ble), and you find GTG codons at cor­re­spond­ing po­si­tions in the En­ter­obac­ter cloa­cae ATCC 13047 and Lel­lot­tia sp. PFL01 genome se­quences *). In ad­di­tion, you eas­ily spot the com­plete con­ser­va­tion of the en­tire N‑ter­mi­­nal se­quence up to the an­no­tated ATG start codon. Thus, a brief look at the DNA se­quence con­firms that these two seem­ingly "aber­rant" DnaA pro­teins are in fact "com­plete". If you had looked at >500 DnaA se­quences – as I did, and it wasn't bor­ing at all – you would have been fa­mi­liar with the oc­cur­rence of rare start codons in >30% of the dnaA genes from bac­te­ria across all phyla. But there's a catch. As a DnaA ex­pert, you would prob­a­bly also know that even small de­le­tions in the N‑terminus ren­der DnaA pro­teins in­ca­pable of dimer­iza­tion via their N‑termini (and thus in­ef­fec­tive for repli­ca­tion ini­ti­a­tion, in­hibitory even). So, you have to check the DNA se­quence for com­plete­ness in those cases where the au­to­matic an­no­ta­tion looks flawed. I don't have to add that this caveat does not only ap­ply to dnaA genes. The re­sults of Hecht et al. sug­gest, in ad­di­tion, that al­go­rithm-dri­ven de­tec­tion of start codons will not make the man­ual cu­ration of se­quence data – in some rare cases even ex­per­i­ments, say, pro­tein se­quenc­ing of N‑ter­mini – ob­so­lete any­time soon.

*) De­spite hav­ing a GTG start codon, the three mR­NAs are trans­lated with an N‑terminal Met residue be­cause fMet-tRNAfMet rec­og­nizes this codon. The "V" in­di­cated in Fig­ure 3 as the N‑terminal residue of the E. can­ce­ro­ge­nus  DnaA pro­tein is for­mally cor­rect but Met would be found when se­quenc­ing the protein's N‑terminus, just as was ex­per­i­men­tally found for E. coli  DnaA whose dnaA gene also has a GTG start codon (un­pub­lished).

 

Other Posts

  • Viruses and the Tree of Life

    by Vin­cent Racaniello — Are viruses liv­ing en­ti­ties? I don't be­lieve so, but it's an en­gag­ing ques­tion for de­bate. On a re­cent episode of 'This Week in Vi­rol­ogy' we con­cluded that it's some­what of a fu­tile ar­gu­ment be­cause every­one has their own view; per­haps our time is bet­ter spent study­ing viruses than...

  • Salmonella's Ex­clu­sive In­testi­nal Restau­rant

    by Elio — In­ter­est­ing, how we get car­ried away by ex­cit­ing con­cepts. I am think­ing about how the study of pathogens has fo­cused so much on the mi­crobes' vir­u­lence fac­tors, by which we mean nasty sub­stances such as tox­ins, ad­hesins, and in­vasins that par­tic­i­pate di­rectly in the dis­ease process. Their study has dom­i­nated mi­cro­bial…

  • Some Like it Hot

    by S. Mar­vin Fried­man — How can ther­mophilic bac­te­ria not only sur­vive, but ac­tu­ally pro­lif­er­ate, at el­e­vated tem­per­a­tures that would be lethal to all other forms of life? Af­ter ex­ten­sive re­search dur­ing the past five decades, this ques­tion has been an­swered in a gen­eral way, but the mol­e­c­u­lar ba­sis for this un­usual ca­pa­bil­ity has not been…

  • It Was The Worst of Times, It Was the Best of Times

    by Elio — At the end of the Per­mian pe­riod, about 250 mil­lion years ago and not long be­fore the di­nosaurs ap­peared, life on Earth ex­pe­ri­enced its great­est cat­a­stro­phe: a mass ex­tinc­tion that did away with the vast ma­jor­ity of life forms on land and sea. The ques­tion arises, who ate the car­casses of the de­ceased?

  • Putting Re­dun­dancy to Work

    by Ka­t­rina Nguyen — Ge­netic re­dun­dancy, when two genes en­code for the same func­tion, is wide­spread among many or­gan­isms. Re­dun­dant genes con­fer an ad­van­tage: if one gene is lost, its part­ner can sub­sti­tute and the phe­no­type of the or­gan­ism will not change. It is un­clear how re­dun­dancy is main­tained dur­ing evo­lu­tion since se­lec­tion against…

  • Strep­to­myces stacks Z‑rings...

    Pic­tures Con­sid­ered #52 by Christoph — 𝘚𝘵𝘳𝘦𝘱𝘵𝘰𝘮𝘺𝘤𝘦𝘴 stacks Z‑rings... into lad­ders in­side its aeri­al/spo­ro­gen­ic hy­phae. If this sen­tence strikes you as a bar­rage of jar­gon or just gib­ber­ish, don't des­pair, help is com­ing. First, en­joy the "light show" in the 20‑se­cond video clip on the right side. It runs in loop, so you can fol­low more than...