Of Terms in Bi­ol­ogy: Metage­nomic Bin­ning

by Christoph

Summer's al­most gone. Imag­ine you're strolling along the shores of a lake en­joy­ing nature's col­ors dur­ing sun­set. Sparkle catches your eyes where the lake lan­guidly laps against the shore. You start pon­der­ing whether mi­crobes — and if so which ones, and how many dif­fer­ent — cause these glis­ten­ing, some­what slimy foam flakes at the shore. Sure enough, you take sam­ples! A bit of mud, some spoons of wa­ter, and sep­a­rately a lit­tle of that slimy foam. Home again, your mi­cro­scope re­veals the splen­did plen­i­tude of small things but con­sid­er­ing that only a neg­li­gi­ble per­cent­age is cul­ti­vat­able, you pass the sam­ples to the se­quenc­ing guys of your de­part­ment. They had just un­boxed their next-gen­er­a­tion se­quenc­ing ma­chines and avidly get them run­ning with your sam­ples.

A cou­ple of days later you get no­ti­fied by the se­quencers that they col­lected ~80 gi­ga­bytes of data for you, rep­re­sent­ing roughly 10,000,000 con­tigs with an av­er­age length of 4 kb, as­sem­bled from the se­quence reads. Dur­ing sam­ple prepa­ra­tion, they had first fil­tered out all larger pro­tists and sub­se­quently re­moved the smaller virus-size par­ti­cles. So far, they had only processed your 'mud sam­ple' but this pile of se­quences could ac­count — by a back-of-the-en­ve­lope cal­cu­la­tion — for 2,000 bac­te­r­ial genomes, each with a 10-fold cov­er­age, but cer­tainly with a heavy over-rep­re­sen­ta­tion of the more abun­dant or­gan­isms. Whack!

Now 'metage­nomic bin­ning' en­ters the stage. Ob­vi­ously, an­a­lyz­ing and sort­ing your bulky data set calls for highly ef­fi­cient soft­ware (= al­go­rithms + pa­ra­me­ter set­tings + data in­put + 'start' but­ton) that does not par­a­lyze the com­put­ing fa­cil­i­ties of your in­sti­tute for days. And you bet­ter for­get about your lap­top. Metage­nomics is, ac­cord­ing to a de­f­i­n­i­tion by Chen &Pachter, "the ap­pli­ca­tion of mod­ern ge­nomics tech­niques to the study of com­mu­ni­ties of mi­cro­bial or­gan­isms di­rectly in their nat­ural en­vi­ron­ments, by­pass­ing the need for iso­la­tion and lab cul­ti­va­tion of in­di­vid­ual species." The ini­tial step in an­a­lyz­ing the se­quence data in or­der to get the tax­o­nomic di­ver­sity pro­file of your en­vi­ron­men­tal sam­ple is re­ferred to as 'bin­ning'. As a com­pu­ta­tional process, bin­ning is con­cep­tu­ally anal­o­gous to es­tab­lished ma­chine learn­ing tech­niques as it in­volves clas­si­fy­ing and/or clus­ter­ing the se­quence con­tigs into spe­cific bins.

Es­tab­lished bin­ning meth­ods can be clas­si­fied as ei­ther 'guided' or 'un­guided', which refers to whether they are 'trained' by a pre-as­sem­bled set of known se­quences or not. Such 'train­ing sets' can be col­lec­tions of rRNA se­quences — presently con­sid­ered the 'gold stan­dard' to as­sess mi­cro­bial di­ver­sity in an en­vi­ron­men­tal sam­ple — or genes rep­re­sent­ing par­tic­u­lar meta­bolic path­ways. Bin­ning meth­ods can also be guided by ref­er­ence data rep­re­sent­ing the var­i­ous bac­te­r­ial taxa ('tax­on­omy-de­pen­dent') or be de­lib­er­ately 'tax­on­omy in­de­pen­dent', which is par­tic­u­larly ap­peal­ing for not al­ways "round­ing up the usual sus­pects." Fi­nally, ex­ist­ing bin­ning meth­ods are — with re­spect to the al­go­rithms ap­plied — ei­ther align­ment-based or com­po­si­tion-based, or they com­bine both ap­proaches. Com­po­si­tional prop­er­ties of DNA se­quences like GC per­cent­age and oligonu­cleotide us­age pat­terns — re­flect­ing codon us­age in the mostly com­pact bac­te­r­ial genomes — can be math­e­mat­i­cally vec­tor­ized. Re­sults from align­ment-free bin­ning pro­ce­dures are best de­scribed as higher‑dimensional vec­to­r­ial spaces that can be com­pared to known se­quences or mod­els present in ref­er­ence data­bases. Align­ment-free meth­ods are faster and re­quire lesser com­pu­ta­tional re­sources as com­pared to align­ment-based meth­ods, but they re­quire longer in­put se­quences — the ~4 kb con­tigs of your 'mud sam­ple' are fine here.

Fig­ure 1. Emer­gent self-or­ga­niz­ing map (ESOM) show­ing the clus­ter­ing and bin­ning of de novo as­sem­bled metage­nomic data. Each point rep­re­sents a frag­ment of an as­sem­bled scaf­fold. Dark lines be­tween clus­ters show de­fin­i­tive sep­a­ra­tion of genome bins. Col­ors des­ig­nate the genome bin for each scaf­fold frag­ment, the stip­pled white line de­marks the bor­ders of the 'unit map'. Source

Now, sup­pos­ing you chose an align­ment-free bin­ning pro­ce­dure, how can you rep­re­sent the re­sults as a graph­i­cal im­age? For your own pre­vi­ous work, you prob­a­bly re­lied on al­ready 'clas­si­cal' meth­ods for align­ment-based sim­i­lar­ity searches, for ex­am­ple BLAST. Briefly, pair­wise com­par­isons of a set of aligned se­quences are used to con­struct a dis­tance ma­trix by clus­ter analy­sis, for ex­am­ple by CustalW. To gen­er­ate the now fa­mil­iar and in­tu­itively in­ter­preted graphic im­age, the dis­tance ma­trix can then be con­verted into a bi­fur­cat­ing, rooted or un­rooted phy­lo­ge­netic tree by group­ing the most closely re­lated pairs of se­quences. Since re­sults from align­ment-free bin­ning can­not be re­duced to a bi­fur­cat­ing tree, other routes are nec­es­sary to ob­tain a two-di­men­sional pic­ture that is easy to 'read' by hu­mans. By bor­row­ing from in­for­mati­cians work­ing on ar­ti­fi­cial neu­ronal net­works, bioin­for­mati­cians came up with an al­go­rith­mi­cally guided pro­ce­dure that 're­duces' the higher-di­men­sional vec­to­r­ial space to a two-di­men­sional Emer­gent Self-Or­ga­niz­ing Map (ESOM). Note in the ex­am­ple ESOM shown (Fig. 1) the dif­fer­ent qual­i­ties of the 'bor­der re­gions' of the var­i­ous patches of grouped scaf­folds (= grouped over­lap­ping con­tigs) that could be as­signed to dif­fer­ent bac­te­ria: sharp in most cases but at times rather fuzzy. This demon­strates a qual­ity of these maps re­gret­fully miss­ing from con­ven­tional trees where it is al­most im­pos­si­ble to in­di­cate quan­ti­ta­tively nested patches of se­quence sim­i­lar­i­ties due to shared genes among more dis­tantly re­lated or­gan­isms and di­ver­gent se­quence stretches be­tween closely re­lated ones. As an an­a­lyt­i­cal tool, such ES­OMs can be used, for ex­am­ple, to aid in su­per­vis­ing the re­con­struc­tion of var­i­ous com­plete genomes from your highly com­plex en­vi­ron­men­tal 'mud sam­ple'.

You may come up with the un­com­fort­able no­tion that hand­ing over your sam­ples to the se­quenc­ing guys was ac­tu­ally a poor trade: swap­ping a com­plex mi­cro­bial com­mu­nity for a data set just as com­plex. You start won­der­ing whether the novel tech­nique of sin­gle-cell se­quenc­ing could over­come the prob­lem of choos­ing the ap­pro­pri­ate bin­ning method for your data set. Not so! This fas­ci­nat­ing tech­nique is ex­ceed­ingly prone to all sorts of DNA con­t­a­m­i­na­tions, and prepar­ing sin­gle cells is — prob­a­bly an un­der­state­ment — de­mand­ing (more on sin­gle-cell se­quenc­ing here and here). Also, genome re­con­struc­tions by sin­gle-cell se­quenc­ing re­quire ~100 "iden­ti­cal" cells (in paren­the­ses be­cause it de­pends on the thresh­old you set for iden­tity) to achieve full cov­er­age. Hard to come by in most cases, or even im­pos­si­ble as in the case of two fa­vorites of this blog, the mealy­bug en­dosym­bionts: spe­cial­ized ab­dom­i­nal cells (bac­te­ri­o­cytes) of the in­sect host, Planococ­cus citri, har­bor cells of the β‑proteobacterium Trem­blaya prin­ceps, which in turn har­bor cells of the γ‑proteobacterium Moranella en­do­bia. Sure enough, both bac­te­r­ial genomes were se­quenced by the shot­gun-se­quenc­ing ap­proach from par­tially pu­ri­fied in­sect tis­sue. Bin­ning of the se­quence con­tigs was not thus com­pli­cated here but for more com­plex mi­croor­gan­is­mal con­sor­tia the sin­gle-cell se­quenc­ing ap­proach is less at­trac­tive. We can safely ex­pect, how­ever, that sin­gle-cell se­quenc­ing and shot­gun-se­quenc­ing will neatly com­ple­ment each other in the long term.

All the above did not an­swer your ques­tion which mi­crobes — and how many dif­fer­ent — caused those glis­ten­ing, some­what slimy foam flakes at the lakeshore. Hope­fully, though, you have gained an ini­tial al­beit su­per­fi­cial un­der­stand­ing of the pow­er­ful tools bioin­for­mati­cians are de­vel­op­ing not only to un­cover but fi­nally to quan­ti­ta­tively un­der­stand the big pic­ture, the plen­i­tude and di­ver­sity of the small things.

 

Other Posts