Fine Read­ing: Ex­plor­ing the Mi­cro­bial Dark Mat­ter

by Merry Youle

Hic sunt dra­cones

We live in a world run by mi­crobes, the vast ma­jor­ity of which we have yet to iden­tify or name. We can only re­fer to them col­lec­tively as the mi­cro­bial dark mat­ter (MDM). How­ever you de­fine a prokary­otic species, and how­ever you tally them once iden­ti­fied, there is a huge gap be­tween the 12,000 or so validly-named species and the to­tal num­ber on our planet, cur­rently es­ti­mated to be in the mil­lions. The only ev­i­dence we have for the ex­is­tence of that un­cul­tured mob is ei­ther a small sub­unit ri­bo­so­mal RNA (SSU rRNA) se­quence or some hazily-clas­si­fi­able metage­nomic reads. As the speed of se­quenc­ing goes up and the cost goes down, this sort of ev­i­dence ac­crues ever more rapidly, fur­ther widen­ing the gap. The chal­lenge at hand is to find out more about the or­gan­isms that make up that dark mat­ter.

Fig­ure 1. The ma­rine mac­ro­bial dark mat­ter. Hic sunt dra­cones. Source

Our Past Bias

What do we know so far? By build­ing trees from SSU rRNA se­quences, we now know that we share the planet with at least 60 prokary­ote phyla. Half of these are dubbed can­di­date phyla and thus they will re­main, by de­cree, un­til some­one cul­ti­vates one of their mem­bers. The other half, the phyla in good stand­ing, have been sam­pled in a highly bi­ased man­ner. More than 88% of mi­cro­bial iso­lates are from only four bac­te­r­ial phyla (Pro­teobac­te­ria, Fir­mi­cutes, Acti­nobac­te­ria, and Bac­teroidetes). Cul­tured mem­bers of the other phyla are get­ting more at­ten­tion now, de­servedly, but this sheds no light on the as yet un­cul­tured MDM.

Fig­ure 2. Sta­tis­ti­cal in­for­ma­tion from the Genomes On­Line Data­base (GOLD) as of Sep­tem­ber 2011. Phy­lo­ge­netic dis­tri­b­u­tion of the 8,448 bac­te­r­ial genome projects. It is time to cut the pie dif­fer­ently. Source

One by One

A re­cent pa­per by Rinke and col­leagues re­ports on work prob­ing this MDM in a painstak­ing way, se­quenc­ing one cell at a time. They do not den­i­grate the value of the com­mu­nity-level in­sights af­forded by metage­nomics; like­wise they ac­knowl­edge that deep se­quenc­ing in low di­ver­sity en­vi­ron­ments has some­times yielded com­plete genome se­quences of the more abun­dant species. How­ever, they put their ef­fort into a project demon­strat­ing the pro­duc­tiv­ity of a com­ple­men­tary ap­proach: sin­gle cell ge­nomics.

In this study they sam­pled nine sites rep­re­sent­ing ma­rine and fresh­wa­ter en­vi­ron­ments, hy­drother­mal vents, sed­i­ment, and a tereph­tha­late-de­grad­ing biore­ac­tor. They col­lected in­di­vid­ual cells with­out an in­ter­ven­ing cul­tur­ing step by sin­gle-cell flow sort­ing, then am­pli­fied their genomes and screened them based on their SSU rRNA genes. This net­ted them a col­lec­tion of 201 genomes of the more abun­dant rep­re­sen­ta­tives of 21 bac­te­r­ial and eight ar­chaeal lin­eages. Se­quenc­ing yielded 201 draft genomes, 40% com­plete on av­er­age. For com­par­i­son, as of Sep­tem­ber 2011, a to­tal of 2907 mi­cro­bial genome projects had been com­pleted. Al­though the sin­gle cell genomes as­sem­bled in this study are not com­plete, they nev­er­the­less make a sig­nif­i­cant jump in the ge­nomic to­tals as well as the di­ver­sity ex­plored.

Fig­ure 3. A his­tor­i­cal tally of com­pleted mi­cro­bial genome projects recorded in the Genomes On­Line Data­base (GOLD). Source

The Re­wards ?

What in­sights were gained? With the first de­tailed ge­nomic data on many of the can­di­date phyla now in hand, the re­searchers could con­firm their lo­ca­tion within the Tree of Life and their re­la­tion­ships to other phyla. For ex­am­ple, a new phy­lum was added to the Planctomycetes–Verrucomicrobia–Chlamydiae (PVC) su­per­phy­lum. Sim­i­larly, they pro­posed a new su­per­phy­lum to ac­com­mo­date the Nanoar­chaeota (fea­tured on this blog here and here) and four can­di­date ar­chaeal phyla that also have small genomes and very small cell size.

Genes of known func­tion pro­vided a few glimpses of the meta­bolic ca­pa­bil­i­ties as­so­ci­ated with var­i­ous novel lin­eages: hy­dro­gen me­tab­o­lism is wide­spread; genes for car­bon fix­a­tion are com­mon among the Ar­chaea. In nu­mer­ous in­stances genes or path­ways pre­vi­ously as­so­ci­ated with only one do­main were found to be present in a novel lin­eage of an­other do­main. First prize here goes to a Nanoar­chaeon with a gene from the slime mold Dic­tyostelium, the first known in­stance of hor­i­zon­tal gene trans­fer from a eu­kary­tote to an Ar­chaeon. The list of bac­te­r­ial genes ac­quired by Ar­chaea has also length­ened to in­clude com­plete bac­te­r­ial sigma fac­tors and multi-do­main alar­mones. Al­though the Ar­chaeota don't make pep­ti­do­gly­can, two Nanoar­chaeota are ap­par­ently mak­ing some use of a bac­te­r­ial lytic murein trans­g­ly­co­sy­lase, per­haps as a weapon or to fa­cil­i­tate friendly cell-cell in­ter­ac­tions with bac­te­ria. The list goes on and on.

In­clud­ing these genomes in their analy­ses of 893 pub­licly avail­able metagenomes pro­vided new or im­proved clas­si­fi­ca­tion of 340 mil­lion reads. Is 340 mil­lion a lot or a lit­tle? It is less than a per­cent on av­er­age, more than 2% of the to­tal in 19 metagenomes, and up to 20% in metagenomes from the same en­vi­ron­ments as were sam­pled in this study.

The gains in se­quenc­ing method­olo­gies come with a cost. It is one thing to ac­cu­mu­late se­quence data, quite an­other to re­veal the ac­tiv­i­ties of the mi­cro­bial world. The bot­tle­neck is shift­ing from data ac­qui­si­tion to analy­sis. In re­sponse has come the de­vel­op­ment of 'clus­ter­ing' ap­proaches that al­low com­puter analy­ses within a rea­son­able amount of time. Ex­ten­sive meta­data about sam­pled en­vi­ron­ments adds com­plex­ity to the stor­age and in­ter­pre­ta­tion of metage­nomic data. Mean­while, a grow­ing num­ber of con­served hy­po­thet­i­cal pro­teins of un­known func­tion await ge­netic and phys­i­o­log­i­cal in­ves­ti­ga­tion.

Syn­tro­phy

Metage­nomics, ge­nomics, culturing—each of these feeds the oth­ers. Sin­gle cell genomes en­able phy­lo­ge­netic clas­si­fi­ca­tion of pre­vi­ously un­clas­si­fied metage­nomic reads. Metage­nomics re­veals pop­u­la­tion struc­ture, di­ver­sity, and the meta­bolic ca­pa­bil­i­ties of the com­mu­nity as a whole. Cul­tur­ing is a pre­req­ui­site for as­sign­ing spe­cific func­tions to the gene se­quences gen­er­ated by the other two. Much re­mains to be done on all fronts. The ma­jor­ity of the mi­cro­bial metage­nomic reads still can't be clas­si­fied be­yond the do­main level. Key com­mu­nity play­ers may be over­looked be­cause they have not been cul­tured. And the mi­cro­bial dark mat­ter still of­fers a vast terra incog­nita beck­on­ing to ea­ger ex­plor­ers.

 

Ref­er­ence

Rinke C, Schwien­tek P, Sczyrba A, Ivanova NN, An­der­son IJ, Cheng JF, Dar­ling A, Mal­fatti S, Swan BK, Gies EA, Dodsworth JA, Hed­lund BP, Tsi­amis G, Siev­ert SM, Liu WT, Eisen JA, Hal­lam SJ, Kyr­pi­des NC, Stepanauskas R, Ru­bin EM, Hugen­holtz P, Woyke T (2013). In­sights into the phy­logeny and cod­ing po­ten­tial of mi­cro­bial dark mat­ter. Na­ture, 499 (7459), 431−437. PMID 23851394

 

Other Posts

  • Lac operon: Wait… what?! Noise in the base­ment?

    by Christoph — If you've ever come across the topic of 'gene reg­u­la­tion in bac­te­ria' in your bi­ol­ogy classes at high school, col­lege or uni­ver­sity, you've in­evitably learned about a clas­sic: the lac operon of E. coli. Well, this it's not an exam here! Just ask your­self what you re­mem­ber with­out the help of an AI.

  • Phage DNA: Go­ing with the Flow

    by Merry Youle — We've heard it said so of­ten, it must be true. Af­ter tailed phages ad­sorb to their host and bind se­curely to their spe­cific re­cep­tor on the cell sur­face, they in­ject their DNA into the cell and the in­fec­tion is off and run­ning. This in­jec­tion no­tion arose nat­u­rally from the clas­sic…

  • Welkin John­son, As­so­ciate Blog­ger

    In my first year of grad­u­ate school I learned that retro­viruses are unique among an­i­mal viruses in hav­ing a "fos­sil" record, com­prised of the hun­dreds of thou­sands of en­doge­nous proviruses em­bed­ded in the genomes of vir­tu­ally all an­i­mal species (in­clud­ing hu­mans).  Those num­bers sound big, but these rem­nants of an­cient epi­demics are an un­der­es­ti­mate, pos­si­bly...

  • Pro­tein Origami

    by Merry Youle — Pri­ons are pro­teins of ill re­pute, as they keep com­pany with "mad cow dis­ease," Alzheimer's, and other neuro­degenerative mal­adies. But pri­ons are not for­eign pathogens; they are al­ter­na­tive states of nor­mal cel­lu­lar pro­teins. Pro­teins be­hav­ing as pri­ons typ­i­cally as­sume a char­ac­ter­is­tic ß‑sheet sec­ondary struc­ture and ag­gre­gate to form ex­cep­tion­ally sta­ble fibers called amy­loids. Small amy­loid ag­gre­gates can...

1 Comment
Oldest
Newest Most Voted
12 years ago

Great ar­ti­cle — re­ally fas­ci­nat­ing, es­pe­cially the preva­lence of shared genes among seem­ingly dis­tant rel­a­tives. I'm cu­ri­ous — is di­rec­tion of trans­fer eas­ily in­ferred? e.g. from Dic­tyostelium to the Nanoar­chaeon?
Merry replies: Eas­ily? Maybe. Some­times. If the trans­fer is re­cent, the trans­ferred genes will still dis­play the codon us­age and GC pat­tern of the source genome which could be dis­tinctly dif­fer­ent from that in the re­cip­i­ent. An­other tac­tic would be to look for the genes in rel­a­tives of the pu­ta­tive source and re­cip­i­ent or­gan­isms. One would pre­dict that the trans­ferred genes would be widely dis­trib­uted through the source genus or fam­ily or even a higher tax­o­nomic group­ing, but lim­ited to one re­cip­i­ent and its close rel­a­tives. There likely are other ways one could use....but those two come to mind first.