Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluation of the Biolog MicroStation system for yeast identification

One hundred and fifty-nine isolates representing 16 genera and 53 species of yeasts were processed with the Biolog MicroStation System for yeast identification. Thirteen genera and 38 species were included in the Biolog database. For these 129 isolates, correct identifications to the species level were 13.2, 39.5 and 48.8% after 24, 48 and 72 hours incubation at 30 degrees C, respectively. Three genera and 15 species which were not included in the Biolog database were also tested. Of the 30 isolates studied, 16.7, 53.3 and 56.7% of the isolates were given incorrect names from the system's database after 24,48 and 72 h incubation at 30 degrees C, respectively. The remaining isolates of this group were not identified.

NASA Program Environmental Health

Evaluation of Automated Yeast Identification System

One hundred and nine teleomorphic and anamorphic yeast isolates representing approximately 30 taxa were used to evaluate the accuracy of the Biolog yeast identification system. Isolates derived from nomenclatural types, environmental, and clinica isolates of known identity were tested in the Biolog system. Of the isolates tested, 81 were in the Biolog database. The system correctly identified 40, incorrectly identified 29, and was unable to identify 12. Of the 28 isolates not in the database, 18 were given names, whereas 10 were not. The Biolog yeast identification system is inadequate for the identification of yeasts originating from the environment during space program activities.

McGinnis, M. R.

ToF-SIMS spectral analysis of Shewanella oneidensis MR-1 biofilms

Analysis of bacterial biofilms is particularly challenging and important with diverse applications from systems biology to biotechnology. Among the variety of techniques that have been applied, time-of-flight secondary ion mass spectrometry (ToF-SIMS) has many powerful features in studying the surface characteristics of biofilms. ToF-SIMS offers high spatial resolution, mass resolution, and mass accuracy, which permit surface sensitive analysis of biofilm components. Thus, ToF-SIMS provides a powerful solution to addressing the challenge of bacterial biofilm analysis. This dataset covers ToF-SIMS analysis of Shewanella oneidensis MR-1 isolated from freshwater lake sediment in New York state. The MR-1 strain is known to have metal and sulfur reducing properties and it can be used for bioremediation and wastewater treatment. There is a current need to identify small molecules and fragments produced from bacterial biofilms, especially those from extracellular polymeric substance (EPS). Static ToF-SIMS spectra of MR-1 were obtained using an IONTOF TOF.SIMS V instrument equipped with a 25 keV Bi$^+_3$ metal ion gun. Identified molecules and molecular fragments are compared against known biological databases and the reported peaks have at least 65 ppm mass accuracy. These molecules range from lipids, fatty acids, flavonoids, and quinolones to other naturally occurring organic compounds. It is anticipated that the mass spectral identification of key peaks will assist detection of metabolites, EPS molecules like polysaccharides, and biologically relevant small organic molecules using ToF-SIMS in future surface and interface research.

59 BASIC BIOLOGICAL SCIENCES

ToF-SIMS spectral data analysis of Paenibacillus sp. 300A biofilms and planktonic cells

Analysis of bacterial biofilms is particularly challenging and important with diverse applications from systems biology to biotechnology. Among the variety of techniques that have been applied, time-of-flight secondary ion mass spectrometry (ToF-SIMS) has many promising features in studying the surface characteristics of biofilms. ToF-SIMS offers high spatial resolution and high mass accuracy, which permit surface sensitive analysis of biofilm components. Thus, ToF-SIMS provides a powerful solution to addressing the challenge of bacterial biofilm analysis. This dataset covers ToF-SIMS analysis of Paenibacillus sp. 300A (300A) isolated from the Hanford site in Richland, WA. The strain is known to have metal and sulfur reducing properties and can be used for bioremediation, wastewater treatment, bioengineering and technology development. There is a current need to identify small molecules and fragments produced from bacterial biofilms. Static ToF-SIMS spectra of 300A were obtained using an IONTOF TOF-SIMS V instrument equipped with a 25 keV Bi 3 + metal ion gun. Identified molecules and molecular fragments are compared against known biological databases and the reported peaks have at least 65 ppm mass accuracy. These molecules range from lipids and fatty acids to flavonoids, quinolones, and other naturally occurring organic compounds. It is anticipated that the spectral identification of key peaks will assist detection of metabolites, extracellular polymeric substance molecules like polysaccharides, and biologically relevant small molecules using ToF-SIMS in future surface and interface research of bacterial biofilms.

Biofilms

DiMER

SAND2025-04145O DiMER is a Python based tool that helps researchers understand the functions of genes by searching through multiple biological databases. It takes user-provided data and scans various databases to find the best matches for gene functions, generating a clear summary of results. DiMER identifies the most relevant functional annotations and improves upon previous annotations by replacing instances of "unknown protein function" with more accurate descriptions. DiMER requires minimal setup. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Mageeney, Catherine [Sandia National Lab. (SNL-CA)

Defining criteria for broadly neutralizing HIV antibodies

Over the course of a few years, a small percentage of individuals with HIV-1 develop broadly neutralizing antibodies (bnAbs) capable of neutralizing diverse viruses. Although hundreds of antibodies with neutralizing activity against heterologous viruses have been referred to as bnAbs, there is no universally accepted numerical definition of a bnAb. Here, we will review important elements of HIV neutralizing antibodies and proposed definitions of bnAbs, as well as introduce a web-based tool, CAByN (Choose Antibodies by Neutralization), allowing users to identify antibodies meeting their numerical definitions of a bnAb from data in the Los Alamos HIV Databases CATNAP (Compile, Analyze and Tally NAb Panels) antibody neutralization database. Biological findings from use of CAByN are also presented here, including differential neutralizing activity for certain antibodies across viral clades, and identification of antibodies with suspected incomplete neutralization. Website address: http://hiv.lanl.gov/content/sequence/CABYN/CABYN.html.

59 BASIC BIOLOGICAL SCIENCES

Dose, LET, time and strain dependence of radiation-induced 53BP1 foci in 15 mouse strains ex vivo and associations to in vivo radiation susceptibility

We present a comparative analysis on the repair of radiation-induced DNA damage ex vivo in 15 strains of mice, including 5 inbred reference strains and 10 collaborative-cross strains, of both sexes. Non-immortalized primary skin fibroblasts derived from 76 mice were subjected to both low- and high-LET radiation (0.1, 1 and 4 Gy of X rays; 1.1 and 3 particles/100μm2 of 350 MeV/n 40Ar and 600 MeV/n 56Fe). Automated image quantification of 53BP1 radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET. Similarly to what we had previously reported for immortalized human cell lines [1], we observed a saturation of RIF number with dose at 4h post-irradiation, with more RIF/Gy for lower LET (X rays and 40Ar) compared to 56Fe. However at later time points (24h and above), the trend was inverted with more RIF/Gy for higher LET. Our data suggest that multiple DSBs cluster into RIF: as the linear density of DSBs increases with LET, so does the probability of having more DSBs per RIF, which makes it more difficult for cells to fully resolve high-LET-induced RIF, explaining the hypersensitivity to high-LET radiation despite a low number of RIF. Taking into account the amount of clustering at a given dose and LET, but also the kinetics of DNA damage repair, we introduced a novel mathematical formalism to evaluate the number of remaining RIF over time. We showed that the newly introduced kinetic metrics can be used as surrogate biomarkers for in vivo radiation toxicity, with potential applications in radiotherapy and human space exploration. In particular, we observed an association between the repairable fraction of RIF measured in vitro and survival levels of immune cells collected from irradiated mice. Moreover, the speed of DNA damage repair correlated with spontaneous cancer incidence data collected from the Mouse Tumor Biology database, suggesting a relationship between the efficiency of DSB repair after irradiation and cancer risk. In addition to the efficacy of repair and persistent RIF levels, even the amount of spontaneous foci without irradiation was shown to be strain dependent, indicating that these phenotypes are at least partially driven by genetics, and supporting their potential as indicators of individual radiation sensitivity. [1] Neumaier, T., et al., PNAS, 2012 (8) 109:443

Radiation, DNA damage, repair kinetics

NASA biological and physical sciences databases: who’s the FAIRest of them all?

Conceptual models are a key part of the foundation of scientific study. Scientific data discovery and retrieval are often inaccurate and incomplete because these models are not sufficiently well-incorporated into data retrieval systems. Systems often don’t provide the necessary tools to those producing scientific data to fully and unambiguously annotate them and the result is consumers of the data cannot find them efficiently. The capability of data archives to provide these tools to link data to underlying conceptual models is one of dimensions of the recently developed “FAIR” principles (https://www.go-fair.org/fair-principles/ ), and is key to many automated processes being able to operate on these data, particularly analytics involving artificial intelligence. We used an open-source web service to measure the FAIR compliance of the three data archives operated by NASA for the biological and physical sciences: the Life Sciences Data Archive, the Physical Sciences Informatics database, and GeneLab. The service ingests references to data sets in these archives, and then executes domain-non-specific examinations of these data and metadata that test compliance to the FAIR principles. Of the 22 metrics tested, GeneLab passed 11 (50%), and PSI and LSDA each passed 7 (32%). These data were gathered using only one representative data set from each archive and we anticipate variability in results as we continue to apply these metrics to other data. A preliminary study of the failure traces for each metric suggests there is a wide range of effort and complexity in the enhancements required for each system to elevate FAIR compliance, and this is the subject of continued investigation. This information has been and will likely continue to be important information in planning these enhancements, with the goal of increased readiness of the data for automated processes.

database

VIZARD: analysis of Affymetrix Arabidopsis GeneChip data

SUMMARY: The Affymetrix GeneChip Arabidopsis genome array has proved to be a very powerful tool for the analysis of gene expression in Arabidopsis thaliana, the most commonly studied plant model organism. VIZARD is a Java program created at the University of California, Berkeley, to facilitate analysis of Arabidopsis GeneChip data. It includes several integrated tools for filtering, sorting, clustering and visualization of gene expression data as well as tools for the discovery of regulatory motifs in upstream sequences. VIZARD also includes annotation and upstream sequence databases for the majority of genes represented on the Affymetrix Arabidopsis GeneChip array. AVAILABILITY: VIZARD is available free of charge for educational, research, and not-for-profit purposes, and can be downloaded at http://www.anm.f2s.com/research/vizard/ CONTACT: moseyko@uclink4.berkeley.edu.

Non-NASA Center

A Proxy Method to Bridge LCA Data Gaps Using Automated Material Classification and Probabilistic Under-Specification

Life cycle assessments (LCAs) are essential for understanding the environmental impacts of material production. However, gaps in life cycle inventory (LCI) data for material and chemical inputs present a key challenge for LCA practitioners, especially in the early design stages. Strategies for filling in these gaps require additional time and expertise, which can hinder the LCA’s completion. This study combined automatic material classification and probabilistic under-specification to create a time-efficient method to fill material LCI data gaps. To illustrate the proposed method, proxy environmental impact distributions were generated using publicly available material LCI data classified into the ChemOnt chemical taxonomy using the open-source chemical classification software ClassyFire. Input materials with data gaps were then classified into the same taxonomy, where proxy environmental impact values could be selected from the available distributions to quickly fill in any data gaps. Although these methods were applied to classify material production processes available in the Federal LCA Commons and Ecoinvent databases, they can be applied to any LCA database. This study shows that classifying materials by their chemical structure produces taxonomies with increased granularity relative to industrial classification, improving the ability of under-specified proxy data to be used for differentiating the environmental impacts of competing designs.

biological databases

BEAST DB: Grand-Canonical Database of Electrocatalyst Properties

We present BEAST DB, an open-source database comprised of ab initio electrochemical data computed using grand-canonical density functional theory in implicit solvent at consistent calculation parameters. The database contains over 20,000 surface calculations and covers a broad set of heterogeneous catalyst materials and electrochemical reactions. Calculations were performed at self-consistent fixed potential as well as constant charge to facilitate comparisons to the computational hydrogen electrode. This article presents common use cases of the database to rationalize trends in catalyst activity, screen catalyst material spaces, understand elementary mechanistic steps, analyze the electronic structure, and train machine learning models to predict higher fidelity properties. Users can interact graphically with the database by querying for individual calculations to gain a granular understanding of reaction steps or by querying for an entire reaction pathway on a given material using an interactive reaction pathway tool. BEAST DB will be periodically updated, with planned future updates to include advanced electronic structure data, surface speciation studies, and greater reaction coverage.

database

Dara: Automated Multiple-Hypothesis Phase Identification and Refinement from Powder X-ray Diffraction

Powder X-ray diffraction (XRD) is a foundational technique for characterizing crystalline materials. However, the reliable interpretation of XRD patterns, particularly in multiphase systems, remains a manual and expertise-demanding task. As a characterization method that only provides structural information, multiple reference phases can often be fit to a single pattern, leading to potential misinterpretation when alternative solutions are overlooked. To ease humans’ efforts and address the challenge, we introduce Dara (data-driven automated Rietveld analysis), a framework designed to automate the robust identification and refinement of multiple phases from powder XRD data. Dara performs an exhaustive tree search over all plausible phase combinations within a given chemical space and validates each hypothesis using the BGMN Rietveld refinement routine. Key features include structural database filtering, automatic clustering of isostructural phases during tree expansion, and peak-matching-based scoring to identify promising phases for refinement. When ambiguity exists, Dara generates multiple hypothesis which can then be decided between by human experts or with further characterization tools. By enhancing the reliability and accuracy of phase identification, Dara enables scalable analysis of realistic complex XRD patterns and provides a foundation for integration into multimodal characterization workflows, moving toward fully self-driving materials discovery.

Biological databases

Autogenerating a Domain-Specific Question-Answering Data Set from a Thermoelectric Materials Database to Enable High-Performing BERT Models

We present a method for autogenerating a large domain-specific question-answering (QA) dataset from a thermoelectric materials database. We show that a small language model, BERT, once fine-tuned on this automatically generated dataset of 99,757 QA pairs about thermoelectric materials, affords better performance in the field of thermoelectric materials compared to a BERT model fine-tuned on the generic English-language QA data set, SQuAD-v2. We further show that mixing the two data sets (ours and SQuAD-v2), which have significantly different syntactic and semantic scopes, allows the BERT model to achieve even better performance. The best-performing BERT model fine-tuned on the mixed data set outperforms the models fine-tuned on the other two data sets by scoring an exact match of 67.93% and an F1 score of 72.29% when evaluated on our test data set. This has important implications as it demonstrates the ability to realize high-performing small language models, with modest computational resources, empowered by domain-specific materials data sets which can be generated according to our method.

biological databases

Locating Undocumented Wells Using Historical Oil and Gas Exploration Maps: A Case Study in Osage County, Oklahoma

Undocumented oil and gas wells lack reliable information about their locations and characteristics, making them difficult to identify. These wells can result in unanticipated delays and costs in the development of nearby surface and subsurface resources, and, if improperly plugged, can cause contamination. This study leverages historical petroleum exploration maps to locate such wells, focusing on Osage County, Oklahoma. Two sets of early 20th century oil and gas exploration maps by the United States Geological Survey were georeferenced and analyzed using a computer vision model to detect well symbols. The locations of detected wells were compared to the location of known wells in the database from the Bureau of Indian Affairs Osage Agency to identify potential undocumented wells. The analysis yielded over 500 potential undocumented wells, with dry holes constituting the largest fraction. Field verification confirmed the presence of some undocumented wells. Comparison with prior work revealed limited overlap, underscoring the complementary value of historical oil and gas maps for locating undocumented wells. This approach demonstrates the utility of integrating historical cartographic resources with modern geospatial and machine learning techniques to improve the identification and management of undocumented wells.

Energy - Petroleum

MARLOWE: An Untargeted Proteomics, Statistical Approach to Taxonomic Classification for Forensics

General proteomics research for fundamental science typically addresses laboratory- or patient-derived samples of known origin and composition. However, in a few research areas, such as environmental proteomics, clinical identification of infectious organisms, archeology, art/cultural history, and forensics, attributing the origin of a protein-containing sample to the organisms that produced it is a central focus. A small number of groups have approached this problem and developed software tools for taxonomic characterization and/or identification using bottom-up proteomics. Most such tools identify peptides via database search, and many rely on organism-specific peptides as markers. Our group recently introduced MARLOWE, a software tool for taxonomic characterization of unknown samples based on de novo peptide identification and signal-erosion-resistant strong peptides, which are shared peptides distributed in a taxonomy-dependent manner. In the current work, we further characterize the utility of MARLOWE using publicly available proteomics data from forensically-relevant samples. MARLOWE characterizes samples based on their protein profile, and returns ranked organism lists of potential contributors and taxonomic scores based on shared strong peptides between organisms. Overall, the correct characterization rate ranges between 44 and 100%, depending on the sample type and data acquisition parameters (with lower numbers associated with lower-quality data sets). MARLOWE demonstrates successful characterization of true contributors and close relatives, and provides sufficient specificity to distinguish certain microbial species. MARLOWE demonstrates its ability to provide insight into potential taxonomic sources for a wide range of sample types without prior assumptions about sample contents. As a result, this approach can find utility in forensic science and also broadly in bioanalytical applications that utilize proteomics approaches for taxonomic characterization.

Bacteria

Database of Nonaqueous Proton-Conducting Materials

This work presents the assembly of 48 papers, representing 74 different compounds and blends, into a machine-readable database of nonaqueous proton-conducting materials. SMILES was used to encode the chemical structures of the molecules, and we tabulated the reported proton conductivity, proton diffusion coefficient, and material composition for a total of 3152 data points. The data spans a broad range of temperatures ranging from -70 to 260 °C. To explore this landscape of nonaqueous proton conductors, DFT was used to calculate the proton affinity of 18 unique proton carriers. The results were then compared to the activation energy derived from fitting experimental data to the Arrhenius equation. It was found that while the widely recognized positive correlation between the activation energy and proton affinity may hold among closely related molecules, this correlation does not necessarily apply across a broader range of molecules. This work serves as an example of the potential analyses that can be conducted using literature data combined with emerging research tools in computation and data science to address specific materials design problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH