Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data lineage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

express: Extensible, high-level workflows for swifter ab initio materials modeling

In this work, we introduce an open-source Julia project, express, an extensible, lightweight, high-throughput, high-level workflow framework that aims to automate ab initio calculations for the materials science community. express is shipped with well-tested workflow templates, including structure optimization, equation of state (EOS) fitting, phonon spectrum (lattice dynamics) calculation, and thermodynamic property calculation in the framework of the quasi-harmonic approximation (QHA). It is designed to be highly modularized so that its components can be reused across various occasions, and customized workflows can be built on top of that. Users can also track the status of workflows in real-time, and rerun failed jobs thanks to the data lineage feature express provides. Finally, two working examples, i.e., all workflows applied to lime and akimotoite, are also presented in the code and this paper.

36 MATERIALS SCIENCE↗

Issues and Solutions for Bringing Heterogeneous Water Cycle Data Sets Together

The water cycle research community has generated many regional to global scale products using data from individual NASA missions or sensors (e.g., TRMM, AMSR-E); multiple ground- and space-based data sources (e.g., Global Precipitation Climatology Project [GPCP] products); and sophisticated data assimilation systems (e.g., Land Data Assimilation Systems [LDAS]). However, it is often difficult to access, explore, merge, analyze, and inter-compare these data in a coherent manner due to issues of data resolution, format, and structure. These difficulties were substantiated at the recent Collaborative Energy and Water Cycle Information Services (CEWIS) Workshop, where members of the NASA Energy and Water cycle Study (NEWS) community gave presentations, provided feedback, and developed scenarios which illustrated the difficulties and techniques for bringing together heterogeneous datasets. This presentation reports on the findings of the workshop, thus defining the problems and challenges of multi-dataset research. In addition, the CEWIS prototype shown at the workshop will be presented to illustrate new technologies that can mitigate data access roadblocks encountered in multi-dataset research, including: (1) Quick and easy search and access of selected NEWS data sets. (2) Multi-parameter data subsetting, manipulation, analysis, and display tools. (3) Access to input and derived water cycle data (data lineage). It is hoped that this presentation will encourage community discussion and feedback on heterogeneous data analysis scenarios, issues, and remedies.

Acker, James↗

Addressing and Presenting Quality of Satellite Data via Web-Based Services

With the recent attention to climate change and proliferation of remote-sensing data utilization, climate model and various environmental monitoring and protection applications have begun to increasingly rely on satellite measurements. Research application users seek good quality satellite data, with uncertainties and biases provided for each data point. However, different communities address remote sensing quality issues rather inconsistently and differently. We describe our attempt to systematically characterize, capture, and provision quality and uncertainty information as it applies to the NASA MODIS Aerosol Optical Depth data product. In particular, we note the semantic differences in quality/bias/uncertainty at the pixel, granule, product, and record levels. We outline various factors contributing to uncertainty or error budget; errors. Web-based science analysis and processing tools allow users to access, analyze, and generate visualizations of data while alleviating users from having directly managing complex data processing operations. These tools provide value by streamlining the data analysis process, but usually shield users from details of the data processing steps, algorithm assumptions, caveats, etc. Correct interpretation of the final analysis requires user understanding of how data has been generated and processed and what potential biases, anomalies, or errors may have been introduced. By providing services that leverage data lineage provenance and domain-expertise, expert systems can be built to aid the user in understanding data sources, processing, and the suitability for use of products generated by the tools. We describe our experiences developing a semantic, provenance-aware, expert-knowledge advisory system applied to NASA Giovanni web-based Earth science data analysis tool as part of the ESTO AIST-funded Multi-sensor Data Synergy Advisor project.

Leptoukh, Gregory↗

Data Quality Challenges for Analysis Ready Data (ARD)

Data quality plays a critical role in research and applications. The Earth Science Information Partners (ESIP) Information Quality Cluster (IQC) defines four aspects of information quality: Science, Product, Stewardship, and Services. The ESIP IQC has become internationally recognized as an authoritative and responsive resource of information and guidance to data producers and distributors on how to implement data quality standards and best practices for their science data systems, datasets, and data/metadata dissemination services. In recent years, cloud computing environments have provided scale-up capabilities such as data archives and services, enabling interdisciplinary science and applications. More value-added products are expected from data service providers, including Analysis Ready Data (ARD). ARD refers to data that has been preprocessed into a form that allows immediate analysis by the end user, processed to a minimum set of requirements and provides interoperability over time and across multiple datasets. Once a dataset has been developed from its original form to produce ARD, what quality characteristics should the derived dataset or ARD possess? Also, is it safe to assume that the quality of the ARD is consistent with the quality of the source data, or are there special attributes to an ARD that would warrant a secondary, independent quality assessment? What provenance (also called “data lineage”) information needs to be included in ARD? It is important to answer these questions, especially given the ease of use of ARD, and the consequent temptation by users to trust ARD without understanding the limitations or possible variations in quality compared to the source data. In this presentation, we will discuss data quality challenges for ARD products and services and introduce IQC for participation.

data quality↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

CovTransformer: A transformer model for SARS-CoV-2 lineage frequency forecasting

With hundreds of SARS-CoV-2 lineages circulating in the global population, there is an ongoing need for predicting and forecasting lineage frequencies and thus identifying rapidly expanding lineages. Accurate prediction would allow for more focused experimental efforts to understand pathogenicity of future dominating lineages and characterize the extent of their immune escape. Here, we first show that the inherent noise and biases in lineage frequency data make a commonly-used regression-based approach unreliable. To address this weakness, we constructed a machine learning model for SARS-CoV-2 lineage frequency forecasting, called CovTransformer, based on the transformer architecture. We designed our model to navigate challenges such as a limited amount of data with high levels of noise and bias. We first trained and tested the model using data from the UK and the USA, and then tested the generalization ability of the model to many other countries and US states. Remarkably, the trained model makes accurate predictions two months into the future with high levels of accuracy both globally (in 31 countries with high levels of sequencing effort) and at the US-state level. Our model performed substantially better than a widely used forecasting tool, the multinomial regression model implemented in Nextstrain, demonstrating its utility in SARS-CoV-2 monitoring. Assuming a newly emerged lineage is identified and assigned, our test using retrospective data shows that our model is able to identify the dominating lineages 7 weeks in advance on average before they became dominant. Overall, our work demonstrates that transformer models represent a promising approach for SARS-CoV-2 forecasting and pandemic monitoring.

60 APPLIED LIFE SCIENCES↗

Genetic Predictive Factors for Nonsusceptible Phenotypes and Multidrug Resistance in Expanded-Spectrum Cephalosporin-Resistant Uropathogenic Escherichia coli from a Multicenter Cohort: Insights into the Phenotypic and Genetic Basis of Coresistance

Antimicrobial resistance in urinary tract infections (UTIs) is a major public health concern. This study aims to characterize the phenotypic and genetic basis of multidrug resistance (MDR) among expanded-spectrum cephalosporin-resistant (ESCR) uropathogenic Escherichia coli (UPEC) causing UTIs in California patient populations. Between February and October 2019, 577 ESCR UPEC isolates were collected from patients at 6 clinical laboratory sites across California. Lineage and antibiotic resistance genes were determined by analysis of whole-genome sequence data. The lineages ST131, ST1193, ST648, and ST69 were predominant, representing 46%, 5.5%, 4.5%, and 4.5% of the collection, respectively. Overall, 527 (91%) isolates had an expanded-spectrum β-lactamase (ESBL) phenotype, with bla CTX-M-15 , bla CTX-M-27 , bla CTX-M-55 , and bla CTX-M-14 being the most prevalent ESBL genes. In the 50 non-ESBL phenotype isolates, 40 (62%) contained bla CMY-2 , which was the predominant plasmid-mediated AmpC (pAmpC) gene. Narrow-spectrum β-lactamases, bla TEM-1B and bla OXA-1 , were also found in 44.9% and 32.1% of isolates, respectively. Among ESCR UPEC isolates, isolates with an ESBL phenotype had a 1.7-times-greater likelihood of being MDR than non-ESBL phenotype isolates (P < 0.001). The cooccurrence of bla CTX-M-15 , bla OXA-1 , and aac(6')-Ib-cr within ESCR UPEC isolates was strongly correlated. Cooccurrence of bla CTX-M-15 , bla OXA-1 , and aac(6')-Ib-cr was associated with an increased risk of nonsusceptibility to piperacillin-tazobactam, cefepime, fluoroquinolones, and amikacin as well as MDR. Multivariate regression revealed the presence of bla CTX-M-55 , bla TEM-1B , and the ST131 genotype as predictors of MDR.

59 BASIC BIOLOGICAL SCIENCES↗

The Impact of Rise of the Andes and Amazon Landscape Evolution on Diversification of Lowland terra-firme Forest Birds

Since the 19th Century, the unmatched biological diversity of Amazonia has stimulated a diverse set of hypotheses accounting for patterns of species diversity and distribution in mega-diverse tropical environments. Unfortunately, the evidence supporting particular hypotheses to date is at best described as ambiguous, and no generalizations have emerged yet, mostly due to the lack of comprehensive comparative phylogeographic studies with thorough trans-Amazonian sampling of lineages. Here we report on spatial and temporal patterns of diversification estimated from mitochondrial gene trees for 31 lineages of birds associated with upland terra-firme forest, the dominant habitat in modern lowland Amazonia. The results confirm the pervasive role of Amazonian rivers as primary barriers separating sister lineages of birds, and a protracted spatio-temporal pattern of diversification, with a gradual reduction of earlier (1st and 2nd) and older (> 2 mya) splits associated with each lineage in an eastward direction. (The easternmost tributaries of the Amazon, the Xingu and Tocantins Rivers, are not associated with any splits older than > 2 mya). For the suboscine passerines, maximum-likelihood estimates of rates of diversification point to an overall constant rate over the past 5 my (up to a significant downturn at 300,000 y ago). This "younging-eastward" pattern may have an abiotic explanation related to landscape evolution. Triggered by a new pulse of Andean uplift, it has been proposed that modern Amazon basin landscapes may have evolved successively eastward, away from the mountain chain, starting approximately 10 mya. This process was likely based on the deposition of vast fluvial sediment masses, known as megafans, that may have extended progressively and in series eastward from Andean sources. This process plausibly explains the progressive extinction of original Pebas wetland of western-central Amazonia by the present fluvial landsurfaces of a more terra-firme type. The youngest landsurfaces thus lie furthest from the mountains. In this scenario major drainages were also reoriented in wholesale fashion away from a northerly orientation generally towards the east and an Atlantic Ocean outlet. The advance of megafans is best seen by the location of axial rivers such as the Orinoco and Mamore which lie against the cratonic margins furthest from the Andes, at the distal ends of major megafan ramparts. More importantly, other major river courses in western-central Amazonia will have been established at progressively younger dates with distance eastward. If this landscape-sequence scenario is accurate, it parallels the progressive younging of the passerine lineages. The bird DNA data appears to confirm strongly the pervasive role of Amazonian rivers--as primary barriers separating sister lineages of birds, and thus probably as facilitaters of bird speciation. We show for the first time that a general spatio-temporal pattern of diversification for terra-firme lineages in the Amazon is associated with rivers ("younging-eastward"), and furthermore parallels a specific scenario of regional drainage evolution.

Aleixo, Alexandre↗

A genomic timescale of prokaryote evolution: insights into the origin of methanogenesis, phototrophy, and the colonization of land

BACKGROUND: The timescale of prokaryote evolution has been difficult to reconstruct because of a limited fossil record and complexities associated with molecular clocks and deep divergences. However, the relatively large number of genome sequences currently available has provided a better opportunity to control for potential biases such as horizontal gene transfer and rate differences among lineages. We assembled a data set of sequences from 32 proteins (approximately 7600 amino acids) common to 72 species and estimated phylogenetic relationships and divergence times with a local clock method. RESULTS: Our phylogenetic results support most of the currently recognized higher-level groupings of prokaryotes. Of particular interest is a well-supported group of three major lineages of eubacteria (Actinobacteria, Deinococcus, and Cyanobacteria) that we call Terrabacteria and associate with an early colonization of land. Divergence time estimates for the major groups of eubacteria are between 2.5-3.2 billion years ago (Ga) while those for archaebacteria are mostly between 3.1-4.1 Ga. The time estimates suggest a Hadean origin of life (prior to 4.1 Ga), an early origin of methanogenesis (3.8-4.1 Ga), an origin of anaerobic methanotrophy after 3.1 Ga, an origin of phototrophy prior to 3.2 Ga, an early colonization of land 2.8-3.1 Ga, and an origin of aerobic methanotrophy 2.5-2.8 Ga. CONCLUSIONS: Our early time estimates for methanogenesis support the consideration of methane, in addition to carbon dioxide, as a greenhouse gas responsible for the early warming of the Earths' surface. Our divergence times for the origin of anaerobic methanotrophy are compatible with highly depleted carbon isotopic values found in rocks dated 2.8-2.6 Ga. An early origin of phototrophy is consistent with the earliest bacterial mats and structures identified as stromatolites, but a 2.6 Ga origin of cyanobacteria suggests that those Archean structures, if biologically produced, were made by anoxygenic photosynthesizers. The resistance to desiccation of Terrabacteria and their elaboration of photoprotective compounds suggests that the common ancestor of this group inhabited land. If true, then oxygenic photosynthesis may owe its origin to terrestrial adaptations.

Methane/metabolism↗

The expression and function of the achaete-scute genes in Tribolium castaneum reveals conservation and variation in neural pattern formation and cell fate specification

The study of achaete-scute (ac/sc) genes has recently become a paradigm to understand the evolution and development of the arthropod nervous system. We describe the identification and characterization of the ac/sc genes in the coleopteran insect species Tribolium castaneum. We have identified two Tribolium ac/sc genes - achaete-scute homolog (Tc-ASH) a proneural gene and asense (Tc-ase) a neural precursor gene that reside in a gene complex. Focusing on the embryonic central nervous system we find that Tc-ASH is expressed in all neural precursors and the proneural clusters from which they segregate. Through RNAi and misexpression studies we show that Tc-ASH is necessary for neural precursor formation in Tribolium and sufficient for neural precursor formation in Drosophila. Comparison of the function of the Drosophila and Tribolium proneural ac/sc genes suggests that in the Drosophila lineage these genes have maintained their ancestral function in neural precursor formation and have acquired a new role in the fate specification of individual neural precursors. Furthermore, we find that Tc-ase is expressed in all neural precursors suggesting an important and conserved role for asense genes in insect nervous system development. Our analysis of the Tribolium ac/sc genes indicates significant plasticity in gene number, expression and function, and implicates these modifications in the evolution of arthropod neural development.

Non-NASA Center↗

Data-Driven Whole-Genome Clustering to Detect Geospatial, Temporal, and Functional Trends in SARS-CoV-2 Evolution

Current methods for defining SARS-CoV-2 lineages ignore the vast majority of the SARS-CoV-2 genome. We develop and apply an exhaustive vector comparison method that directly compares all known SARS-CoV-2 genome sequences to produce novel lineage classifications. We utilize data-driven models that (i) accurately capture the complex interactions across the set of all known SARS-CoV-2 genomes, (ii) scale to leadership-class computing systems, and (iii) enable tracking how such strains evolve geospatially over time. We show that during the height of the original Omicron surge, countries across Europe, Asia, and the Americas had a spatially asynchronous distribution of Omicron sub-strains. Moreover, neighboring countries were often dominated by either different clusters of the same variant or different variants altogether throughout the pandemic. Analyses of this kind may suggest a different pattern of epidemiological risk than was understood from conventional data, as well as produce actionable insights and transform our ability to prepare for and respond to current and future biological threats.

Jacobson, Daniel↗

Mapping the soil microbiome functions shaping wetland methane emissions

Accounting for only 8% of Earth’s land cover, freshwater wetlands remain the foremost contributors to global methane emissions. Yet the microorganisms and processes underlying methane emissions from wetland soils remain poorly understood. Over a five-year period, we surveyed the microbial membership and in situ methane measurements from over 700 samples in one of the most prolific methane-emitting wetlands in the United States. We constructed a catalog of 2,502 metagenome-assembled genomes (MAGs), with more than half of the 70 bacterial and archaeal phyla sampled containing novel lineages. Integration of these data with 133 soil metatranscriptomes provided a genome-resolved view of the biogeochemical specialization and versatility expressed over wetland soil spatial and temporal gradients. Centimeter-scale depth differences best explained patterns of microbial community structure and transcribed functionalities, even more than land cover or temporal information. Moreover, while extended flooding restructured soil redox, this perturbation failed to reconfigure the transcriptional profiles of methane-cycling microorganisms, contrasting with theoretically expected responses to hydrological perturbations. Co-expression analyses, coupled with depth-resolved methane measurements, revealed the metabolisms and trophic structures most predictive of methane hotspots. Mapping the spatiotemporal transcriptional patterns on this compendium of biogeochemically classified soil-derived genomes begins to untangle the microbial carbon, energy, and nutrient processing contributing to wetland methane production.

MAG↗

SOX2-driven enhancer landscape defines the transcriptional architecture of retinogenesis

Retinal neurogenesis is mediated by the coordinated activities of a complex gene regulatory network (GRN) of transcription factors (TFs) in multipotent retinal progenitor cells (RPCs). How this GRN mechanistically guides neural competence remains poorly understood. In this study, we present integrated transcriptional, genetic and genomic analyses to uncover the regulatory mechanisms of SOX2, a key factor in establishing neural identity in RPCs. We show that SOX2 is preferentially enriched in the RPC-specific enhancer landscape associated with essential regulators of retinogenesis. Disruption of SOX2 expression impairs retinogenesis, marked by a selective loss of enhancer activity near genes essential for RPC proliferation and lineage specification. We identified the RPC transcription factor VSX2 as a binding partner for SOX2 and, together, SOX2 and VSX2 co-target a core, retina-specific chromatin repertoire characterized by enhanced TF binding and robust chromatin accessibility. This cooperative binding establishes a shared SOX2-VSX2 transcriptional code that promotes the expression of crucial regulators of neurogenesis while repressing the acquisition of alternative lineage cell fate. Our data illuminate fundamental biological insights on how transcription factors act in concert to drive chromatin-based genetic programs underlying retinal neural identity.

Chromatin↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Diploid-dominant life cycles characterize the early evolution of Fungi

Most of the described species in kingdom Fungi are contained in two phyla, the Ascomycota and the Basidiomycota (subkingdom Dikarya). As a result, our understanding of the biology of the kingdom is heavily influenced by traits observed in Dikarya, such as aerial spore dispersal and life cycles dominated by mitosis of haploid nuclei. We now appreciate that Fungi comprises numerous phylum-level lineages in addition to those of Dikarya, but the phylogeny and genetic characteristics of most of these lineages are poorly understood due to limited genome sampling. Here, we addressed major evolutionary trends in the non-Dikarya fungi by phylogenomic analysis of 69 newly generated draft genome sequences of the zoosporic (flagellated) lineages of true fungi. Our phylogeny indicated five lineages of zoosporic fungi and placed Blastocladiomycota, which has an alternation of haploid and diploid generations, as branching closer to the Dikarya than to the Chytridiomyceta. Our estimates of heterozygosity based on genome sequence data indicate that the zoosporic lineages plus the Zoopagomycota are frequently characterized by diploid-dominant life cycles. We mapped additional traits, such as ancestral cell-cycle regulators, cell-membrane– and cell-wall–associated genes, and the use of the amino acid selenocysteine on the phylogeny and found that these ancestral traits that are shared with Metazoa have been subject to extensive parallel loss across zoosporic lineages. Together, our results indicate a gradual transition in the genetics and cell biology of fungi from their ancestor and caution against assuming that traits measured in Dikarya are typical of other fungal lineages.

59 BASIC BIOLOGICAL SCIENCES↗

Long-term Multimodal Recording Reveals Epigenetic Adaptation Routes in Dormant Breast Cancer Cells

Patients with estrogen receptor–positive breast cancer receive adjuvant endocrine therapies (ET) that delay relapse by targeting clinically undetectable micrometastatic deposits. Yet, up to 50% of patients relapse even decades after surgery through unknown mechanisms likely involving dormancy. To investigate genetic and transcriptional changes underlying tumor awakening, we analyzed late relapse patients and longitudinally profiled a rare cohort treated with long-term neoadjuvant ETs until progression. Next, we developed an in vitro evolutionary study to record the adaptive strategies of individual lineages in unperturbed parallel experiments. Our data demonstrate that ETs induce nongenetic cell state transitions into dormancy in a stochastic subset of cells via epigenetic reprogramming. Single lineages with divergent phenotypes awaken unpredictably in the absence of recurrent genetic alterations. Targeting the dormant epigenome shows promising activity against adapting cancer cells. Overall, this study uncovers the contribution of epigenetic adaptation to the evolution of resistance to ETs.

60 APPLIED LIFE SCIENCES↗

Grass Evolutionary Lineages Can Be Identified Using Hyperspectral Leaf Reflectance

Hyperspectral remote sensing has the potential to map numerous attributes of the Earth’s surface, including spatial patterns of biological diversity. Grasslands are one of the largest biomes on Earth. Accurate mapping of grassland biodiversity relies on spectral discrimination of endmembers of species or plant functional types. We focused on spectral separation of grass lineages that dominate global grassy biomes: Andropogoneae (C 4 ), Chloridoideae (C 4 ), and Pooideae (C 3 ). We examined leaf reflectance spectra (350–2,500 nm) from 43 grass species representing these grass lineages from four representative grassland sites in the Great Plains region of North America. Here, we assessed the utility of leaf reflectance data for classification of grass species into three major lineages and by collection site. Classifications had very high accuracy (94%) that were robust to site differences in species and environment. We also show an information loss using multispectral sensors, that is, classification accuracy of grass lineages using spectral bands provided by current multispectral satellites is much lower (accuracy of 85.2% and 61.3% using Sentinel 2 and Landsat 8 bands, respectively). Our results suggest that hyperspectral data have an exciting potential for mapping grass functional types as informed by phylogeny. Leaf-level hyperspectral separability of grass lineages is consistent with the potential increase in biodiversity and functional information content from the next generation of satellite-based spectrometers.

59 BASIC BIOLOGICAL SCIENCES↗