Mak­ing Genes from Scratch? Prokary­otes not so much...

by Christoph

There is now con­vinc­ing ev­i­dence that yes, Eu­kary­otes make novel genes from scratch (see our pre­vi­ous post ). But what's about the Prokary­otes? Turns out an an­swer to this ques­tion is not so straight­for­ward. Let me try any­way...

Fig­ure 1. The num­ber of genes in com­mon (co­re gen­ome; in green) and to­tal num­ber of ho­mo­lo­gous gene fam­i­lies (pan-genome; in blue) for in­creas­ing num­bers of genomes. Num­bers of gene fam­i­lies were es­ti­mated by per­form­ing 1,000 ran­dom dif­fer­ent in­put or­ders of genom­es. Solid lines cor­re­spond to the av­er­age num­ber of gene fam­i­lies ob­tained by tak­ing into ac­count all per­mu­ta­tions. Dashed lines in­di­cate the stan­dard de­vi­a­tion of the mean. The up­per and lower edges of the blue and green ar­eas cor­re­spond to the max­i­mum and min­i­mum num­bers of gene fam­i­lies, re­spec­tively. For 104 se­quenced genomes, the pan-genome and core genome com­prised 6,867 and 1,791 genes, re­spectively. The core genome rep­re­sented 60% of the av­er­age num­ber of genes per genome and 26% of the pan-genome (suppl. Fig­ure 3). Source

When the first two prokary­otic genomes were se­quenced – those of the archeaon Methanococ­cus jan­naschii and the bac­terium Haemophilus in­fluen­zae, both pub­lished in 1995 – mi­cro­bi­ol­o­gists were star­tled by the high num­ber of an­no­tated genes whose prod­ucts were pre­dicted to be pro­teins of un­known func­tion. Ever since, for every newly se­quenced genome one can safely ex­pect about 1/3 of its genes to fall into this cat­e­gory of pre­vi­ously com­pletely un­known genes (an­no­tated as "hy­po­thet­i­cal pro­tein" ). An­other ~1/3 of the genes be­long to a cat­e­gory of genes with­out known func­tion but with ho­mologs in oth­er se­quenced genomes (an­no­tated as "con­served hy­po­the­ti­cal pro­tein" ). And ~1/3 be­long to the "known" genes, that is, those for which ho­mol­o­gous pro­teins have been stud­ied bio­chem­i­cally (rRNA and tRNA genes be­long to this cat­e­gory al­though not cod­ing for pro­teins ). Not al­ways but of­ten this holds true even for genome se­quen­ces of dif­fer­ent isolates/strains from a sin­gle bac­te­r­ial spe­cies! This proved to be an­other big sur­prise, which led mi­crobiologists to de­velop the con­cept of the pan-ge­no­me (see here in STC ). In a nut­shell, the pan-genome is the sum of the 'core genome' for a species which con­tains the genes shared by al­most all iso­lates, the house­keeping genes, plus a plethora of 'ac­ces­sory', or 'vari­able genes' that are found only in one or a few iso­lates. Dig­ging for dif­fer­ences in the 'vari­able genomes' is a pro­mis­ing way to de­tect, for ex­am­ple, tis­sue-tar­geted vir­u­lence genes in path­o­genic strains of bac­te­ria such as Lis­te­ria mono­cy­to­genes (Fig­ure 1). In the case of E. coli the 'vari­able genes' cur­rently out­num­ber the 'core genome' ge­nes by more than a fac­tor of 10, and a fin­ish line is not in sight. But are all these 'vari­able genes' in a species' pan-genome re­ally novel genes, made from scratch?

Most cer­tainly not, here is why. It was ob­served early on that bac­te­r­ial tox­ins are of­ten ex­pressed (= pro­duced ) by pathogens from genes of res­i­dent prophages rather than from 'core genome' ge­nes. For ex­am­ple, Shiga toxin by the Stx‑1 and Stx‑2 prophages of E. coli O157:H7, cholera toxin by the CTXφ prophage of Vib­rio cholerae, and diph­te­ria toxin by the B‑type lyso­genic phage of Co­ry­ne­bac­te­rium diph­te­riae. For­mally, genes that are 'added' to the chro­mo­some of a bac­te­ri­um by in­tegration of a tem­per­ate bac­te­rio­phage can be con­sid­ered – from the view­point of the bac­te­ri­um – new genes within its 'flex­i­ble genome' but they are not novel genes. They have their own long evo­lu­tion­ary his­tory, in a vir­tu­ally un­fath­omable suc­ces­sion of var­i­ous tem­po­rary host cells, and by us­ing var­i­ous routes of HGT for mov­ing from host to host (for ex­am­ple by 'auto­trans­duc­tion', see our ear­lier post here ). It is now widely ac­cepted that phages play a key role in 'shap­ing' the flex­i­ble ge­no­mes of bac­te­ria, and not only those of pathogens. But, and here's a catch, also the bac­te­ria 're­spond' to this con­tin­u­ous in­flux of new genes. Some do so by or­ga­niz­ing their ge­nomes such that the greater part of the flex­i­ble genome is con­fined to ded­i­cated 'is­lands.' These 'is­lands' ap­pear to be in­volved in a rapid turnover of genes (that is, gain and loss rates for new ge­nes are bal­anced, slightly bi­ased to­wards loss ).

Fig­ure 2. The num­ber of GOS hits per gene, us­ing Pro­chlorococcus mar­i­nus MIT9301 genes as queries, plot­ted against po­si­tion along the chro­mo­some. Shaded re­gions rep­re­sent ge­nomic is­lands, af­ter Cole­man et al. (2006). Flex­i­ble genes with low rep­re­sen­ta­tion in the (Global Ocean Sur­vey) GOS dataset tend to be lo­cated in ge­no­mic is­lands. The num­ber of GOS hits per gene is nor­mal­ized to gene length and plot­ted as hits per gene, per 1,000 bp. Source

This was found in two stud­ies from Penny Chisholm's lab de­scrib­ing the de­tec­tion of five par­tic­u­larly 'HGT-prone' re­gions in the genomes of sev­eral ma­rine Syne­chococ­cus and twelve Prochloro­coc­cus iso­lates (ad­dressed in an ear­lier post by Merry ), which could be sort-of 'cal­i­brated' against each other, and against metage­nomic data from ocean sam­ples (Fig­ure 2). Their analy­sis in­di­cated that many of these 'flex­i­ble genes' do not fit to the phy­logeny of the 'core genes.' They vary greatly in type and num­ber be­tween the var­i­ous eco­types (= sub­types ) of the iso­la­tes, and, most im­por­tantly, po­ten­tially con­fer an adap­tive ad­van­tage to the spe­cific niche of their 'hosts'. The idea that genes, which po­ten­tially help cells to adapt to a no­vel niche, are not brought along all the way by the 'immi­grants' but are al­ready present in that niche, in the lo­cal vi­rome and ready to be picked up via HGT, is still some­what... un­fa­mil­iar, to put it mildly (maybe less so for hard­core ma­rine mi­cro­bi­ol­o­gists ). The Pro­chlo­rococci (Cyanobac­te­ria) are no ex­cep­tion in ex­pertly 'or­ga­niz­ing' their flex­i­ble genomes be­cause, for ex­am­ple, a com­pa­ra­bly high plas­tic­ity con­fined to ge­nomic 'is­lands' was also found by an­a­lyz­ing inter‑strain genome vari­a­tion in He­li­cobac­ter py­lori (Ep­silon­pro­teobac­te­ria).

All cells can quickly adapt to chang­ing en­vi­ron­men­tal con­di­tions by ad­just­ing the over­all ex­pres­sion pat­tern of their genome (see, for ex­am­ple, Fig. 1 in Van­Bo­ge­len and Nei­d­hardt (1990) for the heat-shock re­sponse in E. coli ). There's no need for new genes here. Prokary­otic cells also adapt – though not as quickly – by gene du­pli­ca­tions, which can ac­tu­ally be seen as a recombination‑de­pen­dent rudi­men­tary form of reg­u­la­tion of gene ex­pres­sion. Riehle et al. cul­ti­vated an E. coli  B strain for 2,000 gen­er­a­tions un­der stan­dard con­di­tions (37°C ), fol­lowed for six sep­a­rate lin­eages by an­other 2,000 gen­er­a­tions un­der stress, at near‑nonpermissive tem­per­a­ture (41.5°C ). Af­ter se­lec­tion, three of the six 'stressed' lin­eages turned out to have ac­quired a du­pli­ca­tion of >24 kbp (kbp = 103 base­pairs ) around and in­clud­ing rpoS, en­cod­ing the sta­tion­ary phase-spe­cific sigma-fac­tor. Co­in­cid­ing with the oc­cur­rence of the du­pli­ca­tion at dif­fer­ent time points dur­ing growth at 41.5°C, the strains gained a 2065% in­crease in fit­ness when com­pared to their an­ces­tors. Ob­vious­ly, gene du­pli­ca­tions are an im­por­tant mech­a­nism for adap­ta­tion but a du­pli­cated gene is cer­tainly not a new gene. One usu­ally thinks of mu­ta­tions in the con­text of adap­ta­tion, and yes, cells adapt – some­times sur­pris­ingly fast – by 'tweak­ing' their genes. They al­low se­lec­tion to pick from the con­tin­u­ously aris­ing ran­dom mu­ta­tions those that in­crease the func­tion­al­ity of af­fected gene prod­ucts or si­lence genes if the prod­ucts are detri­men­tal (mea­sured as the sur­vival rate of the mu­tant prog­eny cells; 'sur­vival' rep­re­sents the cu­mu­lated and of­ten in­cre­men­tal ef­fects of fa­vorable and detri­men­tal mu­ta­tions ). Fi­nally, cells adapt by the ac­qui­si­tion of new genes via HGT (see above ) or by mak­ing novel genes from scratch (see fur­ther be­low ).

Genome size and genome com­pact­ness are, in gen­eral, not so much an is­sue in Eu­kary­otes. But one rea­son why Prokary­otes rely heav­ily on HGT for the ac­qui­si­tion of new genes is their com­pact genome struc­ture, which is ap­par­ently un­der strong se­lec­tive pres­sure (size not so much as we know one-chro­mo­some genomes of free-liv­ing bac­te­ria rang­ing in size from ~1.5 Mbp to >10 Mbp; Mbp = 106 base­pairs ). The genome of Schizosac­cha­romyces pombe (As­comy­cota ), for ex­am­ple, has a size of 12.5 Mbp and en­codes ~4,900 genes. In com­par­i­son, roughly the same num­ber of ge­nes (~4.400 ) are en­coded by the ap­prox­i­mately 3‑fold smaller genome of E. coli K‑12 (4.6 Mbp ). Since the pro­teins of S. pombe  and E. coli  do not dif­fer sig­nif­i­cantly in their size dis­tri­b­u­tion, the non-cod­ing 'in­ter­ge­nic re­gions' must be and are in fact shorter in the E. coli  genome. 1,155 in­ter­ge­nic re­gions are re­ally short with lengths of <50 bp, 1,889 range in length be­tween 50 and 300 bp, just 478 are 300 – 900 bp long, and there are only 33 in­ter­genic re­gions >900 bp (note that the num­bers of in­ter­genic re­gions do not add-up to the num­ber of genes be­cause ~30% of the ORFs over­lap each other, a hall­mark of com­pact genomes ). I men­tion these num­bers only to em­pha­size that in the com­pact E. coli  genome the 'open space', that is, non-func­tional DNA, that could serve as 'play­ground' for mak­ing genes from scratch is fairly lim­ited. And even more lim­ited if one takes into ac­count that most in­ter­genic re­gions are all but 'non-func­tional' as they ac­com­mo­date tran­scription sig­nals (pro­mot­ers, ter­mi­na­tors, tran­scrip­tion fac­tor bind­ing-sites ) or genes for a ple­tho­ra of sRNAs that have come into fo­cus only more re­cently. Thus Prokary­otes are – un­like Eu­kar­yo­tes, see the pre­vi­ous post – very lim­ited in evolv­ing novel genes by ac­cu­mu­lat­ing mu­ta­tional chan­ges in in­ter­genic re­gions. HGT by phages, on the other hand, al­lows the cells to aquire new genes 'en bloc' to adapt to chang­ing en­vi­ron­men­tal con­di­tions.

An anal­ogy pops up that is, al­though an­thro­po­mor­phiz­ing, too tempt­ing to put aside. When you work hard on your old iPad 1 (with its very lim­ited mem­ory space of 16 GB, you know it! ) to de­ve­lop a pro­gram in R, for ex­am­ple, you prob­a­bly find it way more ap­peal­ing to down­load an al­ready ex­is­tent code snip­pet 'from the cloud' rather than cod­ing that part your­self from scratch, in­clud­ing all these un­nerv­ing bug-fix­ing ef­forts. Clearly less ap­peal­ing is the per­spec­tive, how­ever, that when­ever you down­load snip­pets from the web you're in­evitably ex­posed to viruses... But one in­tri­cate way how bac­te­ria even­tu­ally turn phage genes into own new  genes will be ex­plored in the fi­nal se­quel of this »mak­ing genes from scratch« se­ries.

Fron­tispiece: pic­ture by Bud­dhini Sama­ras­inghe, from her blog jar­gonwall.

 

Other Posts

  • Un Tour d'Horizon (1/3)

    by Christoph — When bio­chemists de­ci­phered the ge­netic code in the 1960s (the triplet 'al­pha­bet' for amino acids whose de­fined or­der make up the pro­tein 'words' ) it was, and still is, the most com­pelling evi­dence for a com­mon "de­scent with mod­i­fi­ca­tions" (Charles Dar­win) of all life on Earth: the al­pha­bet is the same in all life forms. This holds true...

  • Your Name and Oc­cu­pa­tion?

    by Jen­nifer Gutier­rez — How would you like to be able to ask this ques­tion of an in­di­vid­ual bac­terium in its nat­ural habi­tat – and get a truth­ful an­swer? It turns out that there is a way to do just that, thanks to a cool and in­no­v­a­tive tech­nique called Sec­ondary Ion Mass Spec­trom­e­try (SIMS). A cur­rent ver­sion, known as...

  • Beget­ting the Eu­karya: An Un­ex­pected Light

    by Franklin M. Harold — Con­cern­ing the ori­gin of eu­kary­otic cells, much has been writ­ten but al­most every­thing re­mains to be set­tled. No one dis­putes that mi­to­chon­dria de­rive from free-liv­ing bac­te­ria that es­tab­lished an in­ti­mate sym­bi­otic re­la­tion­ship with a host of some kind and pro­gres­sively turned into or­ganelles, work­horses of me­tab­o­lism, and a hall­mark of eu­kary­otic…

  • This Sum­mer I Went A‑conferencing

    by Elio — In July, I at­tended the 2016 Gor­don Re­search Con­fer­ence on Mi­cro­bial Stress. This is an au­gust meet­ing with 2 decades to its name. The sub­jects pre­sented were broader than the name of the meet­ing sug­gests and in­cluded many as­pects of mi­cro­bial phys­i­ol­ogy and ge­net­ics, touch­ing even on ecol­ogy and evo­lu­tion. "Stress," es­pe­cially in the world of...