How E. coli Rose to Promi­nence – Again, One More Ad­den­dum (2|2)

by Christoph

In the fourth and, for the time be­ing, last episode of this Ecol­i­ma­nia in STC, I take a look around: Do we know the di­ver­sity of the species E. coli ? And what can we learn from ana­lyses of its pangenome ? Re­minder: Roberto looked into the past of How E. coli's Rose to Promi­nence, plus An Ad­den­dum, and I looked ahead in Just One More Ad­den­dum.
 

Fig­ure 2.1. Cir­cu­lar rep­re­sen­ta­tion of the E. coli NCTC 86 chro­mo­some. From the out­side to in­side, cir­cle 1 shows the sizes in bp. Cir­cle 2 marks the po­si­tions of re­gions of dif­fer­ence (RODs). Cir­cles 3 and 4 show the po­si­tions of cod­ing se­quences (CDSs) tran­scribed in a clock­wise and an­ti­clock­wise di­rec­tion (the po­si­tion of the replica­tion ori­gin oriC is in­di­cated). Cir­cles 5 to 10 show the po­si­tions of E. coli NCTC 86 genes that have ortho­logues (by rec­i­p­ro­cal FASTA analy­sis) in other E. coli strains: K‑12 MG1655 (red), E24377A (ETEC; yel­low), 042 (ETEC; green), Sakai (O157:H7; blue), IAI39 (Extrain­testinal Path­o­genic E. coli; pink), CFT073 (UPEC; or­ange). Cir­cle 11 shows a plot of G+C con­tent. Cir­cle 12 shows a plot of G+C skew. Note that what ap­pear to be graph­i­cal glitches in each cir­cle are lo­cal ge­nomic in­ver­sions. Source. Fron­tispiece: De­tail from Fig­ure 2.4.

This is not about the di­ver­sity of E. coli in all its splen­dor. I will limit my­self to the quanti­ta­tive as­pects of its ge­nomic di­ver­sity. AllThe­Bac­te­ria, the pub­licly ac­ces­si­ble data­base just pub­lished by Hunt et al. (2024) and con­tain­ing 300,000+ E. coli genomes, triples the num­ber of the previ­ous "661k" col­lec­tion. The sheer vol­ume of these "stamp col­lec­tions" should suf­fice to cover the ge­netic di­ver­sity of E. coli within the lim­its of the known sam­pling bias: most genome se­quences were ob­tained from path­o­genic and com­men­sal iso­lates from hu­mans, do­mes­tic mam­­mals and birds.

A note in ad­vance: the 'cir­cu­lar plot' im­ages in this post are mainly il­lus­tra­tive and a look at the over­all pic­ture is suf­ficient, but the large num­ber of data points in each im­age can­not re­ally be grasped. All im­ages rep­re­sent com­promises in re­gard to the reso­lu­tion (in pix­els) that can be achieved in print or the file size for web­pages. Fur­ther­more, it is not yet technical­ly fea­si­ble to re­pro­duce inter­active im­ages on web­sites – and is im­pos­si­ble in print – that would al­low zoom­ing in, which would be cru­cial for a de­tailed eval­u­a­tion of the im­ages.

Do we know the ge­nomic di­ver­sity of the species E. coli ?

In 2017, Dunne et al. (2017) se­quenced the genome of E. coli NCTC 86 and com­pared it to the genomes of six other strains (Fig­ure 2.1). E. coli NCTC 86 is the strain orig­i­nally de­scribed by Theo­dor Es­cherich as Bac­terium coli com­mune, later re­named Es­cherichia coli, and in 1919 de­posited in the NCTC. Its genome has a size of 5.144 Mbp, is clas­si­fied as be­long­ing to phy­lo­­type A (see be­low for "phy­lo­types") and is re­lated to K‑12 strain MG1655 (4.625 Mbp), but is not at or close‑to the base of a (phy­lo­genetic) E. coli species tree (Fig­ure 2.1b). Hence, it is not "the wild type" re­searchers could or should re­fer to.

The cir­cu­lar align­ment of the seven E. coli genomes in Fig­ure 2.1 of­fers an in­struc­tive and ex­em­plary overview of their sim­i­lar­ity and like­wise their di­ver­sity. The se­quences are to a high de­gree co‑linear, re­sult­ing in a syn­tenic arrange­ment of the (col­ored) or­thol­o­gous genes like pearls on strings (in the pa­per, the au­thors do not dis­close the sig­nif­i­cance of the two color shades in each of the 6 cir­cu­lar plots. I as­sume this means dif­fer­ences in se­quence iden­tity). Dun­ne et al. (2017) count 41 re­gions of di­ver­gence (RODs) that co­in­cide in po­si­tion in all ge­nomes. They found dif­fer­ences in gain/loss of sin­gle or a hand­ful of genes that are widely scat­tered through­out the genomes. To­gether, these ROSs and the nu­mer­ous sin­gle gene gain/loss events con­tribute to the ~0.5 Mbp dif­fer­ences in genome size be­tween the seven strains com­pared.

Fig­ure 2.1b. Phy­lo­ge­netic re­la­tion­ships amongst se­quenced E. coli genomes. The genomes of E. al­ber­tii and E. fer­gusonii are in­cluded as out­group se­quences. The tree was ob­tained by max­i­mum like­li­hood analy­sis of a con­cate­nated align­ment of 2173 genes, us­ing the gen­eral time re­versible (GTR/REV) model, with the CAT ap­prox­i­ma­tion of rate het­ero­gene­ity as im­ple­mented in RAxML ver­sion 8.1.3. The num­bers on in­di­vid­ual bran­ches in­di­cate the per­cent­age sup­port from 100 non-para­­met­ric boot­strap repli­cates, per­formed us­ing the rapid al­go­rithm im­ple­mented in RAxML. The branch length in­di­cates the num­ber of sub­sti­tu­tions per site. The ma­jor E. coli phy­lo­ge­netic groups (phy­lo­types) are in­di­cated. filled black cir­cles (●), strains in Fig­ure 2.1. Source

If you look re­ally close at Fig­ure 2.1 you see what ap­pear as glitches in the graph­ics are ac­tu­ally small lo­cal­ized ge­nomic in­ver­sions. Ig­nored in the fig­ure for clar­ity but men­tioned in the text is a large 2.8 Mbp in­ver­sion in the E. coli NCTC 86 genome that is not present in the other six ge­nomes (see here). Intra‑genomic in­ver­sions of any size oc­cur com­par­a­tively fre­quently in the ge­nomes of En­ter­obac­te­ri­aceae, usu­ally co­in­cid­ing with the pre­sence of mul­ti­ple trans­pos­able ele­ments in the genome, and are not re­stricted to dis­tinct re­gions. But since it is in most cases im­possible to deter­mine when in­ver­sions oc­curred dur­ing spe­ci­a­tion of a lin­eage be­cause they can be over­laid by later par­tial over­lap­ping in­ver­sions at any time, their phy­lo­ge­netic sig­nal is fuzzy at best. The large in­ver­sion in the genome of E. coli NCTC 86 can only be clas­si­fied as a late "acqui­sition" be­cause it is not found in any of the other six genomes. (As men­tioned here in STC, the large intra‑genomic in­ver­sion in E. coli K‑12 strain W3110, which stands out when com­pared to the genome of K‑12 strain MG1655, is also such a late "ac­qui­si­tion".)

Ghomi et al. (2024) have re­cently taken a dif­fer­ent ap­proach to tack­ling the ge­nomic diver­sity of E. coli. The au­thors com­pared the es­sen­tial­ity of genes (ORFs) in three E. coli strains with that in strains of closely re­lated species. Of par­tic­u­lar in­ter­est is their re­sult of a ge­nome align­ment that clearly vi­su­al­izes species bound­aries (Fig­ure 2.2). Or­thologs are iden­tifiable in the genomes of other En­ter­obac­te­ri­aceae species for the vast ma­jor­ity of E. coli genes (here ir­re­spec­tive of their es­sen­tial­ity). On av­er­age, how­ever, the or­thol­o­gous genes in the three E. coli genomes (the genome of E. coli BW2113 serv­ing as ref­er­ence) are signi­fi­cantly more sim­i­lar to each other than the cor­re­spond­ing or­thologs in the genomes of ten strains of four other species. This con­clu­sion would not change if the iden­tity classes (100%, >85%, >70%) had been cho­sen dif­fer­ently and with more grada­tions (but would be less eas­ily to vi­su­al­ize in a fig­ure).

Fig­ure 2.2. Genome align­ment for the 12 genomes in this study com­pared to E. coli BW25113 from Keio col­lec­tion, ge­ne­rated with BRIG. The genomes from the in­ner cir­cle to the outer cir­cle are E. coli BW25113, E. coli UPEC ST131 EC958, E. coli UPEC ST131, Sal­mo­nella Typhi­mu­rium A130, Salmo­nella Ty­phimurium D23580, Salmo­nella Ty­phimurium SL3261, Sal­mo­nella Ty­phimurium SL1344, Sal­mo­nella en­ter­i­tidis P125109, Sal­mo­nella Ty­phi Ty2, Cit­robac­ter ro­den­tium ICC168, En­ter­obac­ter cloa­cae NCTC 9394, Kleb­siella pneu­mo­niae Ecl8, and Kleb­siella pneu­mo­niae RH201207, re­spec­tively. (Modi­fied by the au­thor to show a scale (in­ner cir­cle) for E. coli BW25113 and the ap­prox. pos. of oriC in this ge­nome). Source

To get the most out of Fig­ure 2.2, you need to know that al­though Ghomi et al. (2024) call it an "align­ment" in the leg­end, the lay­out of the graphic is some­what more com­pli­cated. It is a pro­jection of an or­tholog from each of the other 12 genomes onto the po­si­tion of this or­tholog in the genome of E. coli BW25113, re­gard­less of the po­si­tion of this or­tholog in its own genome (I wish I could de­scribe it in a less twisted way). This be­comes im­mediately clear when one looks at the co-lin­ear­ity of the genomes, which, al­though al­most com­plete for S. Ty­phimurium LT2 and E. coli BW 25113 ex­cept for an in­ver­sion in the ter­mi­nus re­gion, is largely in dis­ar­ray in the case of Cit­robac­ter ro­den­tium, for ex­am­ple (see Fig­ure 2.2b). And then I see nine "gaps" in the cir­cle for the ref­er­ence strain E. coli BW25113, where the au­thors ap­par­ently took or­thologs from the other E. coli strains as the ref­er­ence (with­out ex­plaining this).

Taken to­gether, the genomes of the three E. coli strains stand out as a ho­mo­ge­neous group and the align­ment vi­su­al­izes species bound­aries on the back­ground of a high num­ber of or­thologous genes (ORFs) in the other En­ter­obac­te­ri­aceae. How­ever, species bound­aries can not be sim­ply quanti­fied by an av­er­ag­ing cal­cu­la­tion on this lim­ited set of genome se­quences, not the least since in­di­vid­ual or­thologs are by not equally con­served.

It has been known for more than a cen­tury that the species E. coli is phe­no­typ­i­cally di­verse, and that a sim­ple di­vi­sion into "com­men­sal" and "path­o­genic" iso­lates is in­suf­fi­cient. Ex­ten­sive at­tempts to dis­tin­guish and dif­fer­en­ti­ate E. coli iso­lates based on mor­pho­log­i­cal or meta­bolic cri­te­ria did not lead to a fully con­sis­tent differentiation/typification scheme; nor did sero­typing, which led to the known serotypes O, K and H, think of E. coli O157:H7.

Over time, phy­lo­typ­ing be­came the method of choice for dif­fer­en­ti­at­ing E. coli iso­lates/strains by their geno­types. The method is based on PCR analy­ses, to­day also as dig­i­tal PCR (dPCR) of a few se­lected genes, and was ex­plored in the ini­tial stud­ies by Cler­mont et al. (2000). The num­ber of E. coli phylo­types has grown from ini­tially six to now eight, and these are suf­fi­cient to cover >90% of the ge­nomic di­ver­sity of the species E. coli (see Fig­ure 2.1b). All phy­lo­types con­tain vastly dif­fer­ent num­bers of "mem­bers", and this is pri­mar­ily due to the (above men­tioned) known sam­pling bias for E. coli, rather than to real dif­fer­ences in their oc­cur­rence in nat­ural habi­tats. And in case you won­der, it is now con­firmed by nu­mer­ous phyloge­netic stud­ies that Shigella, iso­lated by Kiyoshi Shiga in 1898 (see here in STC), be­longs to the species E. coli and is there­fore listed among the E. coli phy­lo­types. Only col­leagues in me­dical micro­bi­ol­ogy want to keep Shigella as a (tax­o­nomic) species be­cause they are not in­clined to aban­don the ter­mi­nol­ogy for es­tab­lished di­ag­nos­tics and ther­a­pies of shigel­losis.

What can we learn from analy­ses of the pangenome of the species E. coli ?

First of all, it must be noted that to date there is no pub­lished pangenome analy­sis on all ~105 E. coli genomes in the above-men­tioned "661k" col­lec­tion. All pub­lished stud­ies were per­formed on a sig­nif­i­cantly smaller sub­sets of genomes. The algo­rithms nec­es­sary for com­putational pan­ge­nome analy­ses have been, and are still be­ing, de­vel­oped and eval­u­ated on sub­sets. A ma­jor rea­son for the lack of com­pre­hen­sive pangenome stud­ies for E. coli is that the nec­es­sary com­puting times on hi‑perfor­mance com­put­ers are hard to come by. The main fo­cus of bioinfor­ma­ti­cians is there­fore to make their al­go­rithms ever faster with­out com­pro­mis­ing on qual­ity. It's not un­like wet-lab mi­cro­bi­ol­o­gists us­ing all the tricks to make their fas­tid­i­ous pets grow with genera­tion times of days rather than weeks or months.

Pangenome analy­ses (or "pan‑genome" as men­tioned here in STC) have in­tro­duced a num­ber of terms that are now in com­mon use: "...the core genome com­prises the genes that are present in (nearly) all genomes an­a­lyzed. Some au­thors con­sider a soft­core genome (>95% oc­cur­rence) to avoid dis­miss­ing gene fa­mi­ies due to sequencing/annotation ar­ti­facts. The shell genome con­sists of the genes shared by the majori­ty of genomes (10–95% oc­cur­rence). The gene fam­i­lies pre­sent in only one genome or <10% oc­currence are de­scribed as ac­ces­sory or cloud genome" (adopted from Wikipedia).

Tan­toso et al. (2022) found, when an­a­lyz­ing 1,324 genomes, that the E. coli pangenome is an "open pangenome", mean­ing that with every genome added the num­ber of cloud genes in­creases with­out reach­ing a plateau while the num­ber of core genes (or soft­core genes) asymp­tot­i­cally flat­tens out to a min­i­mum (see here). All the stud­ies I am aware of found an "open pan­genome" for E. coli. And, ac­cord­ing to cur­rent es­ti­mates, the E. coli cloud genome com­prises ~80,000 genes/gene fam­i­lies, while the core genome com­prises ~2,500–3,000 genes/gene fa­mi­lies.

With 10,667 E. coli genomes an­a­lyzed, Abram et al. (2021) come clos­est to a com­pre­hen­sive quan­ti­ta­tive pangenome study and the (presently) least blurry "snap­shot" of the ge­nomic di­ver­sity of the species. The au­thors ob­tained this sub­set by "clean­ing", that is by re­mov­ing du­pli­cates and in­con­sis­tent "com­plete genomes" from the 12,602 E. coli and Shigella genomes in the then‑available NCBI Ref­Seq col­lec­tion. Since it is vir­tu­ally im­pos­si­ble to align and com­pare this num­ber of genomes one-to-one and sub­se­quently de­ter­mine the de­gree of re­lat­ed­ness by con­ven­tional clus­ter analy­sis, the au­thors re­sorted to align­ment-free tools that are rou­tinely used in metage­nomics (see here for a primer in STC). Briefly: the genomes were con­verted to "sketches" in an ap­proach based on k‑mers (k‑mer size of 21, sam­pling size of 10,000). (Note that it is in prin­ci­ple pos­si­ble to "re­verse en­gi­neer" the orig­i­nal se­quence form such "sketches"). From these "sketches", a dis­tance ma­trix for all 10,667 genomes was cal­cu­lated us­ing Python, with the ac­ces­sion num­bers of the genomes as columns and rows. This ma­trix of was then clus­tered with Mash, us­ing hi­er­ar­chi­cal clus­tering to pro­duce a heatmap of 'Mash dis­tances', which il­lus­trates the pop­u­la­tion struc­ture of these genomes (Fig­ure 2.3).

Fig­ure 2.3. Heatmap rep­re­sen­ta­tion of 10,667 genomes us­ing Mash dis­tances. The color bars at the top of the heatmap iden­tify the phy­logroups as pre­dicted from the analy­sis. The scale to the left of the den­dro­gram cor­re­sponds to the re­sul­tant clus­ter height of the en­tire dataset ob­tained from hclust func­tion in R. The col­ors in the heatmap are based on the pair­wise Mash dis­tance be­tween the genomes. Teal col­ors rep­re­sent sim­i­lar­ity be­tween genomes with the dark­est teal cor­re­spond­ing to iden­ti­cal genomes re­port­ing a Mash dis­tance of 0. Brown col­ors rep­re­sent low ge­netic sim­i­lar­ity per Mash dis­tance, with the dark­est brown indi­cating a max. dis­tance of ∼0.039. Genomes of rel­a­tive me­dian ge­netic sim­i­lar­ity have the light­est color. Source

As in­di­cated in the leg­end to Fig­ure 2.3, col­ors in the heatmap are based on the pair­wise Mash dis­tance be­tween the genomes. Teal col­ors rep­re­sent sim­i­lar­ity be­tween genomes with the dark­est teal cor­re­spond­ing to iden­ti­cal genomes re­port­ing a Mash dis­tance of 0. Brown col­ors rep­re­sent low ge­netic sim­i­lar­ity per Mash dis­tance, with the dark­est brown in­di­cat­ing a max­imum dis­tance of ∼0.039. Since you can­not read-out num­bers from Fig­ure 2.3, here they are: 70% of the 10,667 E. coli genomes be­long to four phy­logroups: B1 (28%), A (21%), B2‑2 (13%), and Shig2 (8%). phy­logroup C is rep­re­sented by 5% and E2(O157) by 7% of the genomes. The re­main­ing 18% are dis­trib­uted un­evenly over eight phy­logroups and in­clude the 9% "un­assign­ab­le" cases. The method­ol­ogy used in this work is thus able to re­pro­duce known phy­logroups, as well as to iden­tify pre­vi­ously un­char­ac­ter­ized E. coli phy­logroups, with an ac­cept­able amount of "bot­tom sed­i­ment" (the au­thors use the terms 'phy­lo­type' and 'phylo­group' inter­changeably).

The au­then­tic­ity of these 14 phy­logroups is sup­ported by three dif­fer­ent lines of ev­i­dence: 1. phy­logroup-spe­cific core genes, 2. a phy­lo­ge­netic tree con­structed with 2613 single‑copy core genes, and 3. dif­fer­ences in the rates of gene gain/loss/duplication. It goes with­out say­ing that the analy­sis of en­tire genomes al­lows a much greater depth of analy­sis than the PCR-based method, which is lim­ited to a few gene seg­ments. Note that Abrams et al. (2021) counted "core genes", i.e. genes that are present in all genomes of a phy­logroup, and not, as Ghomi et al. (2024) "es­sen­tial genes" (see above).

To clas­sify the 95,525 unassem­bled genomes, the au­thors took an in­ter­est­ing ap­proach: rather than av­er­ag­ing the Mash val­ues over a phy­logroup, they de­fined a "medoid" as the genome that has the low­est av­er­age dis­tance to all other genomes in its phy­logroup when us­ing the Mash val­ues ob­tained for the en­tire set of 10,667 as­sem­bled genomes. By us­ing the medoids of the 14 phy­logroups as a proxy to clas­sify the unassem­bled genomes, they found that two-thirds (67%) be­long to phy­logroups A (23%), B1 (15%), C (15%), and E2(O157) (14%). Strains be­long­ing to phy­logroups B1, C, and E2(O157) are of­ten path­o­genic and of in­ter­est to med­ical re­search and epi­demi­ol­ogy, whereas phy­logroup A in­cludes strains fre­quently used in the lab or genet­ically mod­i­fied strains. Re­searchers work­ing with these strains are of­ten not in­ter­ested in fully as­sem­bling the genomes, which could ex­plain the un­even dis­tri­b­u­tion across the phy­logroups com­pared to that of the com­plete genomes. (the re­main­ing ~31,000 unassem­bled genomes were dis­trib­uted at lower per­cent­ages among the re­main­ing phy­logroups or did not meet the cut­off).
 

To come back to the ini­tial ques­tion of whether we know the ge­nomic di­ver­sity of the species E. coli, I should give a two-part an­swer:

yes, as far as the ap­prox­i­mate size of the core genome is con­cerned (note that "core gene" and "es­sen­tial gene" are not syn­ony­mous). There are about 2,500–3,000 core genes/gene fam­i­lies across all E. coli phy­lo­types, but since the func­tion of their gene prod­ucts is only known for a frac­tion of the genes, al­though a ma­jor frac­tion now, it is not pos­si­ble to say ex­actly how much re­dun­dancy there is in the core genome. And in our at­tempts to ex­trap­o­late from ge­nomic to ge­netic di­ver­sity, we still lack knowl­edge of many reg­u­la­tory path­ways and knowl­edge of their re­spec­tive re­dun­dan­cies.

Also a cau­tious yes for the ap­prox­i­mate size of the cloud genome. Es­ti­mates based on the num­ber of genomes avail­able in 2021 sug­gest that there are about 80,000 cloud genes/gene fam­i­lies, with the ten­dency of a steadily slow­ing in­crease with more genomes added to the col­lec­tions. Since E. coli has an open pangenome, we will have to be con­tent with es­ti­mates also in the fu­ture.

Pangenome analy­ses car­ried out in re­cent years have al­ready led to com­pelling re­sults, de­spite the fact that this type of analy­sis is in its in­fancy. Or rather, is in pre-school age, to be fair. The ap­proach of a Mash-based analy­sis of mul­ti­ple genomes cho­sen by Abram et al. (2021) can, if the pa­ra­me­ters are set ac­cord­ingly, re­pro­duce the known E. coli phy­lo­types to a satis­fac­tory ex­tent and is suit­able for iden­ti­fy­ing new phy­lo­types. Since genome se­quenc­ing is now a lab rou­tine, it is not pre­sump­tu­ous to ex­pect that this an­a­lyt­i­cal tech­nique will re­place PCR-based "Cler­mon­Typing" over time, es­pe­cially when the "medoids" of the 14 phylo­groups are em­ployed to re­duce the com­pu­ta­tional de­mand. A Mash-based analy­sis may even help with the genomes of E. coli strains that defy PCR‑based phy­lo­typ­ing. It is not pos­si­ble to pre­dict whether the num­ber of 14 phy­logroups pro­posed by Abrams et al. (2021) will per­sist, whether they will need fur­ther dif­fer­en­ti­a­tion, or whether new ones will be added when genomes of E. coli from pre­vi­ously un­der­rep­re­sented habi­tats are added to the "bas­ket".
 

Fig­ure 2.4. Gene re­la­tion­ships in the E. coli pangenome. (a) An overview of the co-oc­cur­rence; on click (b) avoid­ance net­works, coloured by con­nected com­po­nent. Se­lected con­nected com­po­nents are num­bered as in the Sup­ple­men­tary Data Files. Source

The sheer num­ber of core genes/gene fam­i­lies plus the daunt­ing num­ber of cloud genes/gene fam­i­lies in the E. coli pangenome and our cur­rent (at best) spotty un­der­stand­ing of the func­tion(s) of the lat­ter frac­tion are a strong re­minder of what W. Ford Doolit­tle said in a 1999 pa­per, namely "that all prokary­otic taxa are in essence im­pre­cisely bounded and ephemeral. We might thus re­al­is­ti­cally look at all prokary­otes as one 'global su­per­or­gan­ism' ... di­vided into sub­pop­u­la­tions – within and be­tween which genes are ex­changed at dif­fer­ent fre­quen­cies." (ex­changed by hor­i­zon­tal gene trans­fer (HGT), that is). One of Doolittle's 'sub­pop­u­la­tions' is Es­cher­ichia coli, which, ac­cord­ing to Kostas Kon­stan­ti­nides, is a "se­quence-dis­crete pop­u­la­tion" (95–100% ANI) and should there­fore be con­sid­ered a taxo­nomic species. Doolittle's 'impre­cisely bounded' is re­flected in the huge cloud genome of E. coli, and that this species is 'ephemeral', well, yes, this is what evo­lu­tion is about af­ter all.

By a pangenome analy­sis of a se­lected sub­set of 400 E. coli genome se­quences Hall et al. (2021) re­vealed nu­mer­ous groups of genes that pref­er­en­tially co‑occur to­gether in a ge­nome (Fig­ure 2.4), and rec­i­p­ro­cally, groups of genes that pref­er­en­tially do not co‑occur with other gene groups in a genome (click here for Fig­ure 2.4b). These au­thors have cut first deep trails into this seem­ingly im­pen­e­tra­ble jun­gle: there is struc­ture in the E. coli pan­­genome, shaped by its evo­lu­tion, and it's not sim­ply a vast heap of core and cloud genes/gene fam­i­lies. Not that we al­ready un­der­stand this struc­ture well enough...

 

Do you want to com­ment on this post? We would be happy about it! Please send us an email, or com­ment on Mastodon or Bluesky.

Other Posts

  • It Comes With The Ter­roir

    by Roberto — There is some­thing very spe­cial about wine, at least for my taste. I do, how­ever, ad­mit to the pos­si­bil­ity that many of our read­ers will won­der what's the big deal about the bev­er­age, it's just fer­mented grape juice af­ter all.

  • Of Terms in Bi­ol­ogy: RNA Ther­mome­ter

    by Christoph — Fancy call­ing a ther­mome­ter "cool"? An RNA Ther­mome­ter (RNAT) is, tech­ni­cally speak­ing, sim­i­lar to a Breguet ther­mome­ter − not named af­ter the fa­mous French avi­a­tion pi­o­neer, but af­ter his great-great-grand­fa­ther Abra­ham-Louis (1747–1823) − in that it re­acts to chang­ing tem­per­a­tures with ther­mal ex­pan­sion or con­trac­tion of a strip: a bi­metallic strip in Breguet's de­vice and par­tially dou­ble-stranded RNA in an "RNA Ther­mome­ter".

  • Feed­ing on Plas­tic

    by Sab­rina Fitzger­ald — Most peo­ple are fa­mil­iar with the con­cept of a car­bon foot­print. Some may also know there is such a thing as a wa­ter foot­print. But who has ever heard of a plas­tic foot­print? Well, soon, more and more peo­ple will have.

  • Cul­ti­vat­ing the An­ces­tors (4|4)

    by Christoph — Hav­ing 1 or 2 'As­gard' ar­chaea un­der the mi­cro­scope – af­ter hav­ing cul­ti­vated them with great ef­fort and even more pa­tience – and look­ing them in the face is ex­cit­ing, but a bit un­satisfying when they are cousins. Are they mav­er­icks or rather typ­i­cal for "Lo­kis"? Here are the por­traits of two more dis­tant re­la­tives, both also cousins: Mar­gulis­ar­chaeum pep­ti­dophila HC1 and Flexarchae­um multipro­tru­sionis SC1.

  • Preach­ing to a Prokary­otic Choir

    by Mark O. Mar­tin — At my small lib­eral arts un­der­grad­u­ate in­sti­tu­tion, there is only one mi­cro­bi­ol­ogy course, and it is gen­er­ally taken by se­niors. Here is an im­age of my Spring 2011 stu­dents, and be­low is the logo that the tal­ented artist Kaitlin Reiss came up for our class T‑shirts. One of my frus­tra­tions…