Engineering PapersSearch

SEARCH · Engineering Papers

Results for “bioengineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs

Feeding from the sun—Successes and prospects in bioengineering photosynthesis for food security

There is an urgent need for increased crop productivity to reduce food insecurity and improve sustainability. Photosynthesis converts sunlight energy into carbohydrates, providing the source of nearly all of humanity’s food. Photosynthesis is a key target for improvement, owing to inherent inefficiencies in the biochemical process. Over the last decade of advancements in bioengineering, strategies to increase the efficiency of photosynthesis were tested with proven enhancements to crop yields in field trials. Simple strategies like increasing the content of photosynthetic proteins have reliably increased photosynthesis and productivity in crops, as have more complex strategies such as bypassing photorespiration. While insertion of carbon-concentrating mechanisms into C3 plants remains an engineering challenge, modeling suggests that achieving that would have the greatest gain for crop improvement. This review discusses the many successes in improving photosynthesis achieved over the past decade and quantifies the potential for future engineering targets to increase crop productivity.

Long, Stephen P. [University of Illinois, Urbana,

The tier system: a host development framework for bioengineering

Development of microorganisms into mature bioproduction host strains has typically been a slow and circuitous process, wherein multiple groups apply disparate approaches with minimal coordination over decades. To help organize and streamline host development efforts, we introduce the Tier System for Host Development, a conceptual model and guide for developing microbial hosts that can ultimately lead to a systematic, standardized, less expensive, and more rapid workflow. The Tier System is made up of three Tiers, each consisting of a unique set of strain development Targets, including experimental tools, strain properties, experimental information, and process models. By introducing the Tier System, we hope to improve host development activities through standardization and systematization pertaining to nontraditional chassis organisms.

09 BIOMASS FUELS

Harnessing evolution: leveraging bacterial isoprenoid pathway diversity toward improved bioengineering strategies

Isoprenoids play vital roles in all domains of life, from beta-carotene in bacteria to heme in humans. Two distinct metabolic pathways have evolved to synthesize the critical precursor of all mature isoprenoids: the mevalonate (MEV) and the methylerythritol phosphate (MEP) pathways. Here, we quantify the extensive inter- and intra-genus heterogeneity in the usage of these two pathways with particular emphasis on rare bacteria that encode both, or neither, pathways. Furthermore, MEP intermediates themselves have non-isoprenogenic roles that may underlie evolutionary pressures driving pathway diversification. Understanding isoprenoid biosynthesis in bacteria offers new avenues toward more sustainable engineering of economically relevant molecules in microbes.

Biotechnology and Synthetic Biology

Bioengineered algal lipids enriched in structured medium- and long-chain triacylglycerols, linoleate, and sn -2 palmitate for human milk fat substitutes

Human milk fat (HMF) contains triacylglycerol (TAG) as its primary component, providing over 50% of the calories for infant nutrition, along with structural and bioactive lipids that are important for immune and nervous system development. Palmitic acid, comprising 20-25% of the fatty acid complement of HMF, is predominantly esterified to the sn -2 position on the glycerol backbone. This regiospecific positioning facilitates absorption as 2-palmitoyl-monoacylglycerol after hydrolysis of the fatty acids at sn -1 and sn -2 by gut lipases. Other features of HMF include enrichment in structured medium- and long-chain triglycerides (MLCTs), and variation in the ratio of oleic acid to linoleic acid with maternal diet and geography. We have engineered Auxenochlorella, an oleaginous green alga, for biosynthesis of an MLCT- and sn -2 palmitate-enriched HMF substitute for infant formula, matching the regioisomeric composition and proportions of the most abundant fatty acids in HMF.

Lin, Jon Y-T [University of California, Berkeley;]

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS

Enhancing lipid production in plant cells through automated high-throughput genome engineering and phenotyping

Plant bioengineering is a time-consuming and labor-intensive process with no guarantee of achieving desired traits. Here, we present a fast, automated, scalable, high-throughput pipeline for plant bioengineering (FAST-PB) in maize (Zea mays) and Nicotiana benthamiana. FAST-PB enables genome editing and product characterization by integrating automated biofoundry engineering of callus and protoplast cells with single-cell matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). We first demonstrated that FAST-PB could streamline Golden Gate cloning, with the capacity to construct 96 vectors in parallel. Using FAST-PB in protoplasts, we found that PEG2050 increased transfection efficiency by over 45%. For proof-of-concept, we established a reporter-gene-free method for CRISPR editing and phenotyping via mutation of high chlorophyll fluorescence 136. We show that diverse lipids were enhanced up to 6-fold using CRISPR activation of lipid controlling genes. In callus cells, an automated transformation platform was employed to regenerate plants with enhanced lipid traits through introducing multigene cassettes. Lastly, FAST-PB enabled high-throughput single-cell lipid profiling by integrating MALDI-MS with the biofoundry, protoplast, and callus cells, differentiating engineered and unengineered cells using single-cell lipidomics. Furthermore, these innovations massively increase the throughput of synthetic biology, genome editing, and metabolic engineering and change what is possible using single-cell metabolomics in plants.

59 BASIC BIOLOGICAL SCIENCES

Data for "Enhancing Lipid Production in Plant Cells through Automated High-Throughput Genome Engineering and Phenotyping"

Plant bioengineering is a time-consuming and labor-intensive process with no guarantee of achieving desired traits. Here, we present a fast, automated, scalable, high-throughput pipeline for plant bioengineering (FAST-PB) in maize (Zea mays) and Nicotiana benthamiana. FAST-PB enables genome editing and product characterization by integrating automated biofoundry engineering of callus and protoplast cells with single-cell matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). We first demonstrated that FAST-PB could streamline Golden Gate cloning, with the capacity to construct 96 vectors in parallel. Using FAST-PB in protoplasts, we found that PEG2050 increased transfection efficiency by over 45%. For proof-of-concept, we established a reporter-gene-free method for CRISPR editing and phenotyping via mutation of high chlorophyll fluorescence 136. We show that diverse lipids were enhanced up to 6-fold using CRISPR activation of lipid controlling genes. In callus cells, an automated transformation platform was employed to regenerate plants with enhanced lipid traits through introducing multigene cassettes. Lastly, FAST-PB enabled high-throughput single-cell lipid profiling by integrating MALDI-MS with the biofoundry, protoplast, and callus cells, differentiating engineered and unengineered cells using single-cell lipidomics. These innovations massively increase the throughput of synthetic biology, genome editing, and metabolic engineering and change what is possible using single-cell metabolomics in plants.

AI/ML

The anaerobic fungus Caecomyces churrovis produces H 2 via a non-bifurcating NADH-dependent enzyme complex

ABSTRACT Hydrogenosomes are mitochondria-derived organelles that produce ATP and H 2 to support energy metabolism in anaerobic eukaryotes. H 2 production allows reoxidation of reduced cofactors generated during fermentative metabolism; however, the metabolic mechanisms for H 2 production in anaerobic eukaryotes remains incompletely understood. In particular, it remains unclear whether anaerobic fungi (AF) hydrogenosomes use a ferredoxin-dependent pathway or a distinct mechanism to regenerate NAD(P) + and link electron transfer to H 2 formation. Here, by combining genomic search, proteomic analysis, and enzymology, we reveal the molecular mechanism for H 2 production in the AF strain Caecomyces churrovis . Our enzyme assays on the organelle fraction of C. churrovis revealed the activity of H 2 :NAD + oxidoreductase but not pyruvate:ferredoxin oxidoreductase, which is usually linked to H 2 formation. We identified genes encoding [FeFe] hydrogenase (Hyd) and NADH dehydrogenase subunits E and F (NuoE, NuoF) in C. churrovis , and confirmed their expression in the isolated hydrogenosomal fractions by proteomic analysis. Combining the individually purified enzymes, we found Hyd and NuoEF proteins formed H 2 directly from NADH independently of ferredoxin, functioning as a non-bifurcating NADH-dependent enzyme rather than an electron-bifurcating enzyme. We identified homologs of hydrogenosomal NuoE, NuoF, and Hyd in many other AF, indicating this pathway is commonly shared among the AF. This work demonstrates the existence of a non-bifurcating NADH-dependent enzyme complex in eukaryotes. Moreover, this complex could potentially be exploited as a target for controlling AF H 2 production and altering fungal metabolism. IMPORTANCE H 2 production is a prominent feature of anaerobic energy metabolism, yet our understanding of eukaryotic mechanisms remains limited. Anaerobic fungi (AF) are key decomposers of lignocellulose and contribute to hydrogen flux in anaerobic environments. Although it has been more than 40 years since the H 2 production from AF was first reported, the molecular mechanism for hydrogenosomal H 2 production and redox balance remains unclear. We demonstrate that AF produce H 2 from NADH utilizing a non-bifurcating NADH-dependent enzyme complex rather than an electron-bifurcating, ferredoxin-dependent variant. We show that this enzyme complex is conserved across multiple AF lineages and thus demonstrate the occurrence of a non-bifurcating NADH-dependent enzyme in eukaryotes. This discovery expands our understanding of eukaryotic hydrogenosomal metabolism, reveals a previously unknown strategy for redox balancing, and highlights potential targets for manipulating H 2 production. These insights have broad implications for microbial energy metabolism, anaerobic ecosystems, and bioengineering of H 2 -producing systems.

Zhang, Bo [Department of Chemical Engineering, Uni

Description of a novel extremophile green algae, Chlamydomonas pacifica , and its potential as a biotechnology host

We present the comprehensive characterization of a newly identified microalga, Chlamydomonas pacifica , originally isolated from a soil sample in San Diego, CA, USA. This species showcases remarkable biological versatility, including a broad pH range tolerance (6–11.5), high thermal tolerance (up to 42 °C), and salinity resilience (up to 2 % NaCl). Its amenability to genetic manipulation and sexual reproduction via mating, particularly between the two opposing strains CC-5697 & CC-5699, now publicly available through the Chlamydomonas Resource Center, underscores its potential as a biotechnological chassis. The biological assessment of C. pacifica revealed versatile metabolic capabilities, including diverse nitrogen assimilation capability, motility and phototaxis. Genomic and transcriptomic analyses identified 17,829 genes within a 121 Mb genome, featuring a GC content of 61 %. The codon usage of C. pacifica closely mirrors that of C. reinhardtii , indicating a conserved genetic architecture that supports a trend in codon preference with minor variations. Phylogenetic analyses position C. pacifica within the core-Reinhardtinia clade yet distinct from known Volvocales species. The lipidomic data revealed an abundance of triacylglycerols (TAGs), promising for biofuel applications and lipids for health-related benefits. Our investigation lays the groundwork for exploiting C. pacifica in biotechnological applications, from biofuel generation to synthesizing biodegradable plastics, positioning it as a versatile host for future bioengineering endeavors.

Alkali tolerant

From clutter to clarity: Emergent neural operators via questionnaire metrics

Real-world datasets in chemical engineering and bioengineering processes—such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials—can often be unlabeled or disorganized, rendering the training of existing supervised learning models ineffective at learning the underlying dynamics. To salvage these datasets for decision-making, we first seek to obtain clarity from the cluttered data. Here, we present a framework for developing “structural” generative models, discovering emergent equations, and constructing efficient emulators from scrambled datasets by integrating unsupervised organizational learning techniques (Questionnaires) with advanced deep learning architectures (Deep Hidden Physics Models and Deep Operator Networks). Our approach is demonstrated on two illustrative model systems: (a) a 1D advection–diffusion partial differential equation representing a winding underground pipe and (b) an ensemble of Stuart–Landau oscillators, an agent-based system of coupled ordinary differential equations. In both cases, we successfully reconstruct meaningful spatial, temporal, and parameter embeddings from scrambled data, enabling good predictions of system dynamics. As a result, we highlight the framework’s potential for broader applications, enabling data-driven system identification in fields with inherently disorganized or hidden parameter spaces.

42 ENGINEERING

ToF-SIMS spectral data analysis of Paenibacillus sp. 300A biofilms and planktonic cells

Analysis of bacterial biofilms is particularly challenging and important with diverse applications from systems biology to biotechnology. Among the variety of techniques that have been applied, time-of-flight secondary ion mass spectrometry (ToF-SIMS) has many promising features in studying the surface characteristics of biofilms. ToF-SIMS offers high spatial resolution and high mass accuracy, which permit surface sensitive analysis of biofilm components. Thus, ToF-SIMS provides a powerful solution to addressing the challenge of bacterial biofilm analysis. This dataset covers ToF-SIMS analysis of Paenibacillus sp. 300A (300A) isolated from the Hanford site in Richland, WA. The strain is known to have metal and sulfur reducing properties and can be used for bioremediation, wastewater treatment, bioengineering and technology development. There is a current need to identify small molecules and fragments produced from bacterial biofilms. Static ToF-SIMS spectra of 300A were obtained using an IONTOF TOF-SIMS V instrument equipped with a 25 keV Bi 3 + metal ion gun. Identified molecules and molecular fragments are compared against known biological databases and the reported peaks have at least 65 ppm mass accuracy. These molecules range from lipids and fatty acids to flavonoids, quinolones, and other naturally occurring organic compounds. It is anticipated that the spectral identification of key peaks will assist detection of metabolites, extracellular polymeric substance molecules like polysaccharides, and biologically relevant small molecules using ToF-SIMS in future surface and interface research of bacterial biofilms.

Biofilms

Conditional guide RNA deactivation by mRNA and small molecule triggers in Saccharomyces cerevisiae

CRISPR interference (CRISPRi) technologies have revolutionized bioengineering by providing precise tools for gene expression modulation, enabling targeted gene perturbation and metabolic pathway optimization. Despite these advances, achieving dynamic control over gene expression by CRISPR-based regulation remains a challenge due to its inherently static nature. Utilizing toehold-mediated strand displacement and ligand-responsive ribozymes (aptazymes), this study introduces switchable guide RNAs (gRNAs) that facilitate tunable gene expression mediated by mRNA or small molecule signals. We demonstrate complete silencing of gRNA via strategically designed 5’ or 3’ extensions that impede the gRNA spacer or the dCas9 handle, with subsequent restoration of function through sequestration or cleavage of the obstructive sequence. The resulting toehold-embedded or aptazyme-embedded gRNAs can be deactivated by specific signals, including two full-length translatable mRNAs and two small molecule triggers, thereby lifting CRISPRi repression on targeted genes. This modular approach allows for gRNA-based biocomputing through multi-layer or multi-input genetic logic gates in Saccharomyces cerevisiae . Offering a versatile strategy for post-CRISPR regulation in response to environmental signals or cellular states, this methodology expands the toolkit in eukaryotic systems for reversible control of gene expression.

Aptazyme

Artificial intelligence methods for protein structure and interaction prediction: Recent advances and challenges

Recent advances in artificial intelligence have introduced novel methods for high-accuracy prediction of protein tertiary structures, protein complex structures, and interactions between proteins and other biomolecules, such as small molecules and nucleic acids. Such advancements are accelerating biomedical research and the development of new protein design and bioengineering methods among many other important biotechnology applications. Here, in this review, we outline the recent advances in protein-centric biomolecular structure and interaction prediction, highlight some major challenges in the field, and discuss potential directions to address them.

Morehead, Alex [Lawrence Berkeley National Laborat

Molecular Modeling and Molecular Dynamics Simulation of a Packed and Intact Bacterial Microcompartment

Bacterial microcompartments (BMCs) are protein-bound organelles found in some bacteria which encapsulate enzymes for enhanced catalytic activity. These compartments spatially sequester enzymes within semipermeable shell proteins and are packed full of enzyme cargoes and metabolites as they fulfill their function. Coupling together recent SAXS and proteomics work, it is possible to develop molecular models for these microcompartments and interrogate enzyme and metabolite dynamics within. Our primary goal of this study is to quantify the permeability of metabolite glyceraldehyde-3-phosphate (G3P) and dihydroxyacetone phosphate (DHAP) across the BMC shell through classical molecular dynamics simulation. The Haliangium ochraceum model of BMC shell (PDB: 6MZX) was used to model an intact BMC of approximately 10 million atoms. Working at this scale presented its own challenges in managing large data sets, with multiple challenges and hardware advances discussed that facilitated this work. Over approximately 750 ns of aggregate simulation, we see multiple permeation events for these metabolites that were added at high concentration through the pores present within BMC shell tiles. When compared to independent permeability estimates for the same metabolites determined through replica exchange umbrella sampling simulations, the permeabilities varied by approximately 3 orders of magnitude. Regardless, the permeability coefficients for both G3P and DHAP are highly similar and very high, such that only very small concentration gradients can be maintained across the BMC shell between the cytosol and BMC interior. The large simulation systems also facilitated comparisons for molecular diffusivity in the crowded environment within the BMC shell. By our estimates, the viscosity within a packed BMC shell is at least 10-fold higher than it would be in neat solution and is the real driver for varying permeability estimates we obtained through simulation. These findings will be used as design inputs for future bioengineering efforts to make products from BMCs, highlighting how permeable BMC shells can be.

Diffusion

Photosynthetic Biohybrid System for Enhanced Abiotic N 2 -to-NH 3 Conversion under Ambient Conditions

Photosynthetic biohybrid systems (PBSs) offer an eco-friendly approach to transforming solar energy into value-added products by integrating biological entities with inorganic semiconductors. However, the chemical conversion capacity of most PBSs has inherent limitations, as whole-cell bacteria and isolated enzymes require fine-tuning of environmental conditions. Here, in this study, we report a new PBS developed by introducing free-standing ceria nanoparticles into the purple membrane (PM) of Halobacterium salinarum archaea, which can unidirectionally transfer charge carriers in response to incident photons, even after separation from living archaea at various conditions. Our microscopy, spectroscopy, and synchrotron X-ray scattering analyses confirm that the electrostatic assembly between ceria and PM creates seamless interfacial contact, thereby enhancing the photocatalytic capacity of ceria. Although the conversion of dinitrogen (N 2 ) to ammonia (NH 3 ) is thermodynamically challenging due to the triple bond in N 2 and a series of charge-transfer reactions, our PM–ceria (PMC) hybrid nanoparticle efficiently produces NH 3 by reducing N 2 using solar energy even under atmospheric pressure and room temperature while simultaneously converting glycerol into value-added derivatives. Additionally, our PMC nanoparticle involves neither toxic/precious metals nor bioengineering processes to achieve enhanced photocatalytic N 2 -to-NH 3 conversion. This study sheds light on the new aspect of PBSs by employing PM to potentially resolve the global energy and environmental challenges posed by the conventional Haber–Bosch process.

Jang, Jinhyeong [Argonne National Laboratory (ANL)

Mechanistic implications of excited high-spin states, spin–spin coupling, and differential [2Fe–2S] + cluster temperature relaxations in the electron-bifurcating NfnABC from Thermococcus sibiricus

Electron bifurcation (EB) is a mechanism of biological energy transduction in which multiple oxidation–reduction (redox) reactions are thermodynamically coupled within a single enzyme, enabling the enzyme to harness the excess free energy from an exergonic process to drive an endergonic process. Because of this unprecedented chemistry, there is interest to translate EB principles to artificial and bioengineered systems, but a hurdle is that knowledge pertaining to the fundamental design principles of EB enzymes remains scarce. Here, we investigated the fundamental physical and electronic properties of electron transfer sites in a spectroscopically uncharacterized member of the BfuABC family of EB enzymes, the NADH-dependent reduced-ferredoxin:NADP + oxidoreductase from Thermococcus sibiricus (Tsi NfnABC). Cryo-EM structures of Tsi NfnABC previously demonstrated that it contains twelve redox cofactors: two flavins (one FAD and one FMN), eight [4Fe–4S] clusters, and two [2Fe–2S] clusters. The FMN, one [4Fe–4S] cluster, and one [2Fe–2S] cluster comprise the bifurcating active site termed the electron-bifurcating flavobicluster (BF-FBC), which is found in all BfuABC family members. By using electron paramagnetic resonance spectroscopy, we identified spectral signatures originating from interactions between the FMN radical and [4Fe–4S] + cluster in the BF-FBC and observed temperature dependent behavior of the BF-FBC's [2Fe–2S] + cluster indicative of moderately slow spin–lattice relaxation. Additionally, we uncovered numerous spectral features corresponding to half-integer, S > ½ spin states of [4Fe–4S] + clusters, including one attributable to the consequences of lysine-ligation of a [4Fe–4S] cluster unique to NfnABC. We contextualize these findings to electron transfer theory and NfnABC's structure. Our insights further the understanding of how enzymes are designed to exert control over electron transfer to conduct thermodynamically challenging reactions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Targeted genetic manipulation and yeast-like evolutionary genomics in the green alga Auxenochlorella

Auxenochlorella spp. are diploid oleaginous green algae whose streamlined genomes can be readily manipulated by homologous recombination, making them highly amenable to discovery research and bioengineering. Vegetatively diploid organisms experience specific evolutionary phenomena, including allodiploid hybridization, mitotic recombination, loss-of-heterozygosity, and aneuploidy; however, studies of these forces have largely focused on yeasts. Here, we present a telomere-to-telomere phased diploid genome assembly of Auxenochlorella UTEX 250-A (haploid length 22 Mb) and introduce a genetic toolkit for site-specific manipulation of the nuclear genome in multiple strains, featuring several selectable markers, inducible promoters, and fluorescent reporters for protein localization. UTEX 250-A is an allodiploid hybrid of Auxenochlorella protothecoides and Auxenochlorella symbiontica, two species differentiated by extensive chromosomal rearrangements. UTEX 250-A haplotypes are a mosaic of each parental species following mitotic recombination, and two chromosomes are trisomic. Loss-of-heterozygosity events are pervasive across Auxenochlorella and can evolve rapidly in the laboratory. High-quality structural annotation yielded ∼7,500 genes per haplotype. Auxenochlorella have experienced gene family loss and reduction, including core photosynthesis genes, and exhibit periodic adenine and cytosine methylation at promoters and gene bodies, respectively. Approximately 10% of genes, especially those involved in DNA repair and sex, overlap antisense long noncoding RNAs, which may participate in a regulatory mechanism. We demonstrate the utility of Auxenochlorella for fundamental research by knockout of a chlorophyll biosynthesis enzyme, and confirm one trisomy by allele-specific transformation. These results demonstrate the generality of several evolutionary forces associated with vegetative diploidy and provide a foundation for the use of Auxenochlorella as a reference organism.

CHL27