Bac­te­r­ial Genomes, IKEA-Style

(Some As­sem­bly Re­quired)

by Janie

What had once seemed dis­crete has turned fuzzy around the edges these days, whether work ver­sus home, or mi­nutes ver­sus days (and pa­ja­mas ver­sus all other alterna­tives). But there's been a sliver of a sil­ver lin­ing in this en masse jum­bling. HaFig­ure 1. A few pieces in the grand puz­zle of li­feforms. Source: Janie Kim. Frontis­piece: DNA con­tain­ing an unna­tural base-pair, X-­Y. Sourceving re­cently taken a class on nu­cleic acid chem­istry and, af­ter look­ing for a way to take my mind off things, now try­ing to learn Es­peranto from Duo­lingo, I've been gen­tly re­minded of how much of a lan­guage the ge­netic code is. Much like "ar­ti­fi­cial" lan­guages like Es­peranto or Tolkien's fif­teen (!) Elvish tongues, ge­nomes, too, are sub­ject to ma­nip­u­la­tion and mod­u­la­tion, by the scientist's hand.

Fig­ure 1. A few pieces in the grand puz­zle of li­feforms. Source: Janie Kim. Frontis­piece: DNA con­tain­ing an unna­tural base-pair, X-­Y. Source

There's been a smat­ter­ing of posts here on the blog that touched on syn­thetic genomes, but the last post that was de­voted to the topic was all the way back from May 2009, not long af­ter Craig Ven­ter an­nounced his first syn­thetic My­coplasma genome. In this piece, Shmuel Razin wrote about Harold Morowitz's vi­sion­ary 1964 pro­posal to syn­the­size a liv­ing cell from its com­po­nents, and then Venter's later an­nounce­ment.

Much more nu­cleic acid tin­ker­ing has tran­spired in the years since 2009. So, here's a look at where the field of build-your-own-genome, bac­te­r­ial edi­tion, stands now as of 2020, with a lit­tle spritz of lin­guis­tics.

Lin­guis­tics and Nu­cleic Acids, From Loan­words to Word­play

The in­ter­tidal zone be­tween nu­cleic acids and hu­man lan­guages bus­tles with analo­gies. There's a lot shared by all those strands of spi­ral­ing chem­i­cal lace­work and the words that roll off our tongues.

Both genomes and hu­man lan­guages pos­sess her­i­ta­ble units and a hi­er­ar­chy of build­ing mean­ing, start­ing from let­ters to codons/words to ge­netic elements/phrases and so on, sub­ject to rules that can be framed math­e­mat­i­cally. In the way a sin­gle in­ser­tion or dele­tion of a nu­cleotide al­ters the func­tion of a gene, a sin­gle comma makes all the dif­fer­ence in mean­ing and mor­bid­ity be­tween "Let's eat, Grandma!" and "Let's eat Grandma!" Gene du­pli­ca­tion lead­ing to func­tional diversifi­ca­tion is not un­like lin­guis­tic redu­pli­ca­tion: in the Maia lan­guage spo­ken in parts of Papua New Gui­nea, 'maia' means 'thing' and its redu­pli­cated form 'ma­ia­maia' came to mean the plural 'things.' In­te­gra­tion of for­eign gene snip­pets through hor­i­zon­tal gene trans­fer, too, has a lin­guis­tic equiva­lent: the Ko­rean word for 'trench coat,' which is pro­nounced buh-buh-ree (버버리), is a loan­word of the Eng­lish brand name Burberry. All those mod­u­lar com­pound words in Ger­man are evoca­tive of NRPS path­ways. Par­al­lels can even be drawn be­tween epi­ge­netic mod­i­fi­ca­tions and lan­guage: the se­quence of let­ters doesn't change, but the con­text that the word is in shapes its mean­ing, as in 'lead' ab­sorp­tion by bac­te­ria could 'lead' to biore­me­di­a­tion. Per­haps even word­play and puns can be thought of as hav­ing a bi­o­log­i­cal anal­ogy: a split be­tween the face-value mean­ing and the mean­ing that is ac­tu­ally in­ter­preted, like mRNA tran­scripts pre- and post-pro­cess­ing, to cap it all off (bad pun #1).

An es­pe­cially fas­ci­nat­ing shared trait: both lan­guages and genomes evolve. Both ex­ist in a state of dy­namic equi­lib­rium be­tween size and com­plex­ity ver­sus ef­fi­ciency. Both grow and shrink with chang­ing cir­cum­stances, whether genome re­duc­tion of an en­dosym­biont or the shrink­ing Eng­lish lex­i­con, dri­ven by the trade-off be­tween size and en­ergy cost.

Fig­ure 2. Photo 51, en­tan­gled in the Frank­lin/­Wat­­son/­Crick saga, kicked gen­etics into high gear in the mid-1900s. The im­age that started it all. Source

Both are also probed by the re­lent­less hu­man cu­rios­ity about the bare bones: what fun­da­men­tal prop­er­ties make them tick? Lin­guists in pur­suit of lin­guis­tic univer­sals and the ever-elu­sive Uni­ver­sal Gram­mar, piec­ing apart lan­guages into their sim­plest units... Sci­en­tists mov­ing from elu­ci­dat­ing the struc­ture of DNA and de­ciphering the ge­netic code all the way to syn­the­siz­ing brand new nu­cle­obases and from-scratch min­i­mal ge­nomes...

How far can this ge­net­ics anal­ogy be taken? The cross­over (bad pun #2) be­tween syn­thetic ge­nomics and Elvish goes pretty far.

Chang­ing the Ge­netic Al­pha­bet: Un­nat­ural Base Pairs

The ge­netic "al­pha­bet" clas­si­cally con­sists of four let­ters and the mod­ern Eng­lish al­pha­bet con­sists of twenty-six. The di­ver­sity that arises from such a small set of nu­cleic acid build­ing blocks is as­tound­ing – but not the end-all. This open-end­ed­ness was em­braced by chemists in the sec­ond half of last cen­tury. By un­der­stand­ing how stan­dard Wat­son-Crick base pair­ing works, could an en­tirely new set of nu­cle­obases be de­signed to func­tion smoothly in repli­ca­tion, tran­scrip­tion, and transla­tion? Per­haps A‑T and C‑G base-pair­ing are not so spe­cial.

Fig­ure 3. "DNA", cour­tesy of xkcd. Source

Re­place Tolkien with a sci­en­tist, the pen with a pipette, and you have a recipe for a bi­o­log­i­cal fan­tasy lan­guage. The field of un­nat­ural base pairs (UBPs) was sparked to life in 1962 when Alexan­der Rich first pro­posed the possi­bility of a third base pair. His idea was picked up years later and ma­te­ri­al­ized by Steve Ben­ner and his group, who shuf­fled around hy­dro­gen bond donors and accep­tors on nu­cle­obases to gen­er­ate new non­canon­i­cal bon­ding pat­terns. UBPs like these could be in­te­grated into DNA and RNA by their re­spec­tive polyme­rases, and the nu­cle­obase al­pha­bet was thereby ex­panded.

But hy­dro­gen bond­ing be­tween bases was not all that it was chalked up to be. As it turned out, poly­merases were happy to work with non-po­lar, hy­dropho­bic analogs of nu­cle­obases. The deter­mining fac­tors for in­cor­po­ra­tion by poly­merases were found to be base-stack­ing and steric fit, and hy­dro­gen bond­ing was knocked off its pedestal. The UBP play­ing field had at that point been un­ex­pectedly ex­panded to in­clude hy­dropho­bic bases.

Fig­ure 4. A sam­pling of nu­cle­obases, both natu­ral and un­nat­ural. Source

More re­cent ad­vances have em­braced both hy­dropho­bic pairs and hy­dro­gen-bond­ing pairs, iron­ing out ini­tial wrin­kles with the orig­i­nal Ben­ner sys­tem. The Romes­berg group sifted through mul­ti­tudes of UBPs to dis­cover that the hy­dropho­bic 5SICS and NaM pair out­per­formed its pre­de­ces­sors in PCR and in vitro tran­scrip­tion. The de­velopment of 'hachi­moji' DNA and RNA came even more re­cently, which added four more syn­thetic nu­cleotides. With hachi­moji, the let­ter count now sits at a pro­vi­sional eight, poised to change as more "xeno-nu­cleic acids" trickle into the lex­i­con.

Ever the ge­net­ics par­al­lel, the Eng­lish A‑to‑Z has also been in flux. An in­ter­est­ing his­tor­i­cal aside: sev­eral let­ters have been re­moved or added since Ye Olde days of Chaucer, a trend vi­su­al­ized in an­tique children's ABC books, which some­times in­cluded the am­per­sand "&" for a to­tal of 27 let­ters (and see here for the his­tory of the "Ye" of "Ye Olde Shoppe" fame, its ori­gins tied with a mix-up in­volv­ing what is now an ob­so­lete let­ter of the alpha­bet). Evo­lu­tion in all its nat­ural and un­nat­ural forms is re­lent­less.

De­signer Genomes

Un­nat­ural bases and ex­pand­ing cod­ing ca­pac­ity com­prise one sub­set of the field of syn­thetic ge­nomics, and cob­bling to­gether whole genomes is an­other. In con­trast to the ex­pan­sion­ist ambi­tions of the UBP world, genome syn­the­sis moved along a more re­duc­tion­ist tra­jec­tory.

Fol­low­ing Khorana's 1972 syn­the­sis of a com­plete gene, the yeast ala­nine tRNA, came rev­o­lu­tions in se­quenc­ing and DNA syn­the­sis tech­nolo­gies. Genome af­ter genome was se­quenced un­til the fo­cus shifted at the turn of the cen­tury, re­dis­trib­ut­ing in­ter­est from se­quenc­ing to­ward synthesi­zing genomes. The year 2002 saw the first syn­thetic genome: the de novo chem­i­cal syn­the­sis of the 7.4‑kb po­liovirus genome. Far be­hind were the days of sky­scraper se­quenc­ing gels and of pointil­lism in ra­dioac­tive black.

In 2008, Craig Ven­ter and his lab un­veiled a syn­thetic My­coplasma genome, all 583-kb worth of spi­ral­ing lad­ders stitched to­gether by a com­bi­na­tion of lig­ases, yeast in­ter­me­di­ates, and forty mil­lion dol­lars. The in­evitable leap from in vitro to in vivo came just two years later. In 2010, Ven­ter an­nounced the first self-repli­cat­ing Franken-cell: a lab-syn­the­sized My­coplasma my­coides genome en­cased in a My­coplasma capri­colum husk. Putting to­gether this 1.08-Mbp genome in­volved splurg­ing on over a thou­sand com­mer­cially man­u­fac­tured se­quences, some con­tain­ing "water­marks" spelling out an email ad­dress, people's names, and quo­ta­tions, to dis­tin­guish the ar­ti­fi­cial genome from the nat­ural one.

While reach­ing mile­stones in whole-genome syn­the­sis, in­ter­est in genome min­i­miza­tion also grew in earnest: what is the ab­solute min­i­mum ge­netic in­for­ma­tion a free-liv­ing cell needs to sur­vive and repli­cate? Thus be­gan the re­duc­tion­ist quest for the bare-bones genome.

In 2016, Ven­ter an­nounced Syn3.0, a My­coplasma with its genome whit­tled down to a mere 473 genes in a to­tal of 531 kbp, the fruit of many rounds of trans­po­son mu­ta­ge­n­e­sis screens. Start­ing with My­coplasma gen­i­tal­ium, the free-liv­ing bac­terium with the small­est genome then known at a mere 525 genes, they had stripped away 10% of its genes. For com­par­i­son, the E. coli genome con­sists of ap­prox­i­mately 4,000 genes in a to­tal of 4.6 Mbp. The minimalist's game of limbo (how min­i­mal can you go?) was well on its way.

But the cost of DNA syn­the­sis is a ma­jor bot­tle­neck to progress in the field. Fab­ri­cat­ing long strands of DNA ex­tend­ing thou­sands of bases de novo is a finicky af­fair, hence, the forty-mil­lion-dol­lar price tag on the Ven­ter My­coplasma. If syn­the­sis costs could be low­ered, progress in whole-genome syn­the­sis would be much faster. To this aim, re­searchers at the ETH Zurich de­vel­oped a com­puter pro­gram to gen­er­ate a syn­the­sis-friendly bac­te­r­ial genome. Their pro­gram per­formed a ma­jor over­haul of the 785-kb "es­sen­tial genome" of Caulobac­ter cres­cen­tus by cut­ting out se­quen­ces that hin­der syn­the­sis, such as repet­i­tive re­gions and re­gions of high GC con­tent, and by speed­ing through the ge­netic the­saurus to swap out over 10,000 bases and in­sert syn­onyms for over 120,000 codons. The pay­off? The Swiss Caulobac­ter re­quired a mere frac­tion of the re­sources of its 2008 My­coplasma pre­de­ces­sor. Ring­ing in at $123,000, the project was a mere 0.3% blip of the Ven­ter cost, and took just one year, as op­posed to twenty, to com­plete.

The size of genomes built from scratch men­tioned so far span from kilo­bases up to the 1.08-Mbp My­coplasma. The year 2019 marked the top­pling of the record. In work pub­lished just one month af­ter the Swiss study, re­searchers in the UK and else­where pro­duced a 4‑Mbp E. coli genome with only 61 codon triplets, a re­sult of 18,214 syn­ony­mous codon sub­sti­tu­tions. To as­sem­ble this sim­plified genome, the group took a page from the syn­thetic chemists' book. In a ret­rosyn­the­sis of sorts, they split their re­coded genome into frag­ments, syn­the­sized these 10-kb pieces, and then as­sem­bled their way back up to the full genome.

The UBP-Genome Merger and a Sight­ing of Xeno­bi­ol­ogy

In the last decade, UBPs and syn­thetic genomes fi­nally met on their in­evitable col­li­sion course. As re­ported in a 2014 pub­li­ca­tion, the un­nat­ural pair d5SICS-dNaM was first in­cor­po­rated into a plas­mid with oth­er­wise nor­mal A‑T and C‑G pairs. This plas­mid was then in­tro­duced into a spe­cial strain of E. coli equipped with an NTP trans­porter from the ma­rine di­atom Phaeo­dacty­lum tri­cor­nu­tum, of bio­fuel fame. When pro­vided with the UBPs ex­oge­nously, the re­sult­ing bac­terium could fun­nel them in through its NTP trans­porter and copy the plas­mid with en­doge­nous machi­nery, and just like that, go to town me­tab­o­liz­ing and jiv­ing and do­ing its repli­cat­ing thing.

Fig­ure 5. Trans­la­tion with an un­nat­ural co­don, an­ti­codon, and amino acid. Source

The next bomb in syn­thetic ge­nomics dropped only a cou­ple years later. Sev­eral labs, in­clud­ing Romesberg's, be­gan cre­at­ing semi-syn­thetic bac­te­ria whose un­nat­ural nu­cle­obases them­selves en­coded un­nat­ural amino acids. All within the con­fines of a liv­ing bac­terium, RNA poly­merase tran­scribed mRNA and tRNA con­tain­ing the UBPs, the ri­bo­some de­coded un­nat­ural codons with the right un­nat­ural an­ti­codons, and the right un­nat­ural pro­teins with un­nat­ural amino acids were pro­duced at the end of the day. Here was a shiny new-fan­gled re­boot of the Cen­tral Dogma.

These semi-syn­thetic or­gan­isms drew a great deal of me­dia at­ten­tion and spec­u­la­tion that the next fore­seeable step might be to drop the "semi." How long would it be un­til the first fully syn­thetic free-liv­ing mi­crobe with a purely UBP genome?

A con­cise an­swer: a long time. The tech­ni­cal is­sues are tricky, as Prof. Ralph Kleiner, a chem­istry pro­fes­sor at Prince­ton Uni­ver­sity who taught the afore­men­tioned nu­cleic acids class, de­scribed: one of the ma­jor hur­dles is coax­ing cells to me­tab­o­lize the un­nat­ural build­ing blocks nec­es­sary for UBP in­stal­la­tion. Romesberg's work with NTP trans­porters rep­re­sents one pos­si­ble so­lu­tion, al­though it is too early to say how gen­eral this strat­egy will be, and it would be far prefer­able to sup­ply sim­pler build­ing blocks like ar­ti­fi­cial nu­cle­o­sides rather than the fully formed nu­cleotide triphos­phate. But even with a re­li­able sys­tem for UBP in­stal­la­tion, in or­der for these pairs to be use­ful, they need to be uniquely rec­og­nized by the pro­tein trans­la­tion ma­chin­ery and ac­ces­sory fac­tors. From en­sur­ing com­pat­i­bil­ity with the ri­bo­some to aminoa­cyl tRNA syn­thetases, as well as codon or­thog­o­nal­ity, there is a lot to align with in in­tra­cel­lu­lar cir­cuitry. Re­mem­ber, Na­ture has had bil­lions of years to fine tune the cur­rent sys­tem. But even with just one pair of un­nat­ural bases, a cell's cod­ing ca­pac­ity jumps from 43 to 63 with a par­al­lel jump in po­ten­tial ap­pli­ca­tions.

As of 2020, the high­est pub­lished num­ber of un­nat­ural codons that can be si­mul­ta­ne­ously deco­ded in a liv­ing cell stands at three, for a to­tal of 67 de­coded codons. There is a slight de­crease in pro­tein yield by the semi-syn­thetic cell, al­beit a rel­a­tively in­signif­i­cant de­crease com­pared to that in cells en­gi­neered with stop codon sup­pres­sion. This high­lights an­other is­sue to over­come: hi­jack­ing the sys­tem of­ten leads to lower ef­fi­ciency, al­though one group is tack­ling this prob­lem by sub­ject­ing E. coli with re­duced genomes to adap­tive lab­o­ra­tory evo­lu­tion.

Some Ex­tra Thoughts, Scat­tered Across the Re­al­ity-to-Sci­ence-Fic­tion Spec­trum

Look­ing past the show­man­ship that seems at times to tinge the field, find­ing min­i­mal genomes and cre­at­ing min­i­mal bac­te­ria will help un­der­stand fun­da­men­tals about what makes life tick. Per­haps DIY-ing bac­te­ria that are just-alive will fur­ther il­lu­mi­nate en­dosym­bionts, their dras­ti­cally re­duced genomes, and their phys­i­ol­ogy. Im­port of nu­cleotide build­ing blocks, for ex­am­ple, is a process shared by the Romes­berg group's d5SICS-dNaM E. coli, en­dosym­bionts, and in­tra­cel­lu­lar par­a­sites. In­deed, the or­gan­isms short­listed for trans­porter-bor­row­ing in­cluded the amoeba en­dosymbiont P. amoe­bophila. To co-opt Richard Feynman's words as syn­thetic bi­ol­ogy co-opts parts of or­gan­isms: "What I can­not cre­ate, I do not un­der­stand." What does it take at the genome level for a cell to for­feit its in­de­pen­dence and take up per­ma­nent res­i­dence within an­other, in what seems a case of mi­cro­bial Stock­holm Syn­drome?

There's also more to gain from the re­cent codon-swap­ping ex­per­i­ments done by the Swiss and UK groups, other than cheaper bac­te­r­ial genome syn­the­sis. En­ter syn­thetic viruses. Man­made viruses, at the core of con­spir­acy the­o­ries run­ning amuck these days, have his­tor­i­cally been pro­duced with a more be­nign mo­tive: de­vel­op­ing vac­cines. For ex­am­ple, swap­ping codons in the po­liovirus ge­nome to run counter to its nor­mal codon bias at­ten­u­ates vir­u­lence, and this cus­tomized inef­fec­tual virus can be used in vac­cines. From po­liovirus to in­fluenza to Ebola, recre­at­ing and re­cod­ing of vi­ral genomes al­lows re­searchers to de­velop and test drugs, vac­cines, and di­ag­nos­tics. This is the same sit­u­a­tion with SARS-CoV‑2, as labs work to recre­ate the virus.

This theme of piec­ing apart and as­sem­bling genome se­quences is rem­i­nis­cent of the new mecha­nical vi­sion of life il­lus­trated in Mem­branes to Mol­e­c­u­lar Ma­chines. The book was pre­vi­ously co­vered on the blog here. Truly, the "cell sim­u­lacrum" has come a long way since the 1970s lipo­somes, those lit­tle ATP-pro­duc­ing glob­ules jig­sawed to­gether from chicken eggs, soy­bean plants, beef heart, and Halobac­te­ria.

Today's semi-syn­thetic mi­crobes could "serve as a plat­form for the cre­ation of new life forms and func­tions," as so starkly worded in a pub­li­ca­tion from the Romes­berg group. But a bac­terium with a com­pletely un­nat­ural al­pha­bet still seems a long way off. It's an idea evoca­tive of world-build­ing in fic­tion, with struc­tures spun into ex­is­tence by imag­i­na­tion and in­vented words. Maybe a fully syn­thetic genome is a pot of gold at the end of the rain­bow, but in the mean­time, we can imag­ine. And we can con­tinue to write spec­u­la­tive fic­tion sto­ries, which all so of­ten slip into re­al­ity years down the line...

 

Other Posts