Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Greedy Sampling and Incremental Surrogate Model-Based Tailoring of Aeroservoelastic Model Database for Flexible Aircraft

This paper presents a data analysis and modeling framework to tailor and develop linear parameter-varying (LPV) aeroservoelastic (ASE) model database for flexible aircrafts in broad 2D flight parameter space. The Kriging surrogate model is constructed using ASE models at a fraction of grid points within the original model database, and then the ASE model at any flight condition can be obtained simply through surrogate model interpolation. The greedy sampling algorithm is developed to select the next sample point that carries the worst relative error between the surrogate model prediction and the benchmark model in the frequency domain among all input-output channels. The process is iterated to incrementally improve surrogate model accuracy till a pre-determined tolerance or iteration budget is met. The methodology is applied to the ASE model database of a flexible aircraft currently being tested at NASA/AFRC for flutter suppression and gust load alleviation. Our studies indicate that the proposed method can reduce the number of models in the original database by 67%. Even so the ASE models obtained through Kriging interpolation match the model in the original database constructed directly from the physics-based tool with the worst relative error far below 1%. The interpolated ASE model exhibits continuously-varying gains along a set of prescribed flight conditions. More importantly, the selected grid points are distributed non-uniformly in the parameter space, a) capturing the distinctly different dynamic behavior and its dependence on flight parameters, and b) reiterating the need and utility for adaptive space sampling techniques for ASE model database compaction. The present framework is directly extendible to high-dimensional flight parameter space, and can be used to guide the ASE model development, model order reduction, robust control synthesis and novel vehicle design of flexible aircraft.

numerical analysi↗

XML technology planning database : lessons learned

A hierarchical Extensible Markup Language(XML) database called XCALIBR (XML Analysis LIBRary) has been developed by Millennium Program to assist in technology investment (ROI) analysis and technology Language Capability the New return on portfolio optimization. The database contains mission requirements and technology capabilities, which are related by use of an XML dictionary. The XML dictionary codifies a standardized taxonomy for space missions, systems, subsystems and technologies. In addition to being used for ROI analysis, the database is being examined for use in project planning, tracking and documentation. During the past year, the database has moved from development into alpha testing. This paper describes the lessons learned during construction and testing of the prototype database and the motivation for moving from an XML taxonomy to a standard XML-based ontology.

prototype databases↗

An infrared spectral database for gas-phase quantitation of volatile per- and polyfluoroalkyl substances (PFAS)

We report the construction of a database of vetted infrared spectra specifically targeting volatile fluorocarbon gases that may be emitted during thermal treatment of per- and polyfluoroalkyl substances (PFAS) to assist understanding of treatment processes and improve quantification. To populate this database, protocols derived from the Pacific Northwest National Laboratory (PNNL) infrared spectral database are used, curtailing the species selection for this data set. Each spectrum in the database is a weighted average derived from 10 or more individual measurements at different partial pressures (static method) or flow rates (dissemination method) to yield good fidelity of both strong and weak infrared signatures, with each composite spectrum ranging from = 6500 cm-1 to = 600 cm-1 with an apodized resolution of 0.112 cm-1. This resolution was chosen to fully resolve all spectral features, recognizing that atmospheric pressure broadening results in nearly all ro-vibrational lines having linewidths = 0.1 cm-1. As an example case, application of the database is demonstrated via identification and quantification of dominant 1H-perfluoroheptane and perfluorohept-1-ene fluorocarbon products resulting from thermal decomposition of perfluorooctanoate (PFOA) below 450 °C.

Infrared, Gas-phase spectra, FTIR, Spectral databa↗

Tripal, a community update after 10 years of supporting open source, standards-based genetic, genomic and breeding databases

Abstract Online, open access databases for biological knowledge serve as central repositories for research communities to store, find and analyze integrated, multi-disciplinary datasets. With increasing volumes, complexity and the need to integrate genomic, transcriptomic, metabolomic, proteomic, phenomic and environmental data, community databases face tremendous challenges in ongoing maintenance, expansion and upgrades. A common infrastructure framework using community standards shared by many databases can reduce development burden, provide interoperability, ensure use of common standards and support long-term sustainability. Tripal is a mature, open source platform built to meet this need. With ongoing improvement since its first release in 2009, Tripal provides full functionality for searching, browsing, loading and curating numerous types of data and is a primary technology powering at least 31 publicly available databases spanning plants, animals and human data, primarily storing genomics, genetics and breeding data. Tripal software development is managed by a shared, inclusive governance structure including both project management and advisory teams. Here, we report on the most important and innovative aspects of Tripal after 11 years development, including integration of diverse types of biological data, successful collaborative projects across member databases, and support for implementing FAIR principles.

59 BASIC BIOLOGICAL SCIENCES↗

The Natural Products Magnetic Resonance Database (NP-MRD) for 2025

The Natural Products Magnetic Resonance Database or NP-MRD (https://np-mrd.org) is a comprehensive, freely accessible, web-based resource for the deposition, distribution, extraction and retrieval of nuclear magnetic resonance (NMR) data on natural products. The NP-MRD was initially established to support compound de-replication and data dissemination for the natural products community. However, that community has now grown to include many users from the metabolomics, microbiomics, foodomics and nutrition science fields. Indeed, since its launch in 2021, the NP-MRD has expanded enormously in size, scope and popularity. The current version of NP-MRD now contains nearly 7X more compounds (281,859 vs. 40,908) and 7X more NMR spectra (5.1 million vs. 817,000) than the first release. More specifically, an additional 4.6 million predicted spectra and another 11,000 spectra simulated from experimental chemical shifts were deposited into the database. Likewise, the number of NMR raw spectral data depositions has grown from a 165 spectra per year to more than 10,000 per year. As a result of this expansion, the number of monthly webpage views has grown from 55 to 20,000 and the number of monthly visitors has increased from 7 to 2500. To address this growth and to better support the expanding needs of its diverse community of users, many additional improvements to the NP-MRD have been made. These include significant enhancements to the data submission process, important improvements to the visualization and display of NMR spectra, notable updates to the database’s spectral search utilities and useful additions to support better NMR spectral analysis/prediction. Significant efforts have also been undertaken to remediate and update many of NP-MRD’s database entries. This manuscript describes these database improvements and expansion efforts, along with how they have been implemented and what future upgrades to the NP-MRD are planned.

Artifical Intelligence↗

SERDP PFAS 2.0 - An infrared spectral database for gas-phase quantitation of volatile per- and polyfluoroalkyl substances (PFAS)

We report the construction of a database of vetted infrared spectra specifically targeting volatile fluorocarbon gases that may be emitted during thermal treatment of per- and polyfluoroalkyl substances (PFAS) to assist understanding of treatment processes and improve quantification. To populate this database, protocols derived from the Pacific Northwest National Laboratory (PNNL) infrared spectral database are used, curtailing the species selection for this data set. Each spectrum in the database is a weighted average derived from 10 or more individual measurements at different partial pressures (static method) or flow rates (dissemination method) to yield good fidelity of both strong and weak infrared signatures, with each composite spectrum ranging from = 6500 cm-1 to = 600 cm-1 with an apodized resolution of 0.112 cm-1. This resolution was chosen to fully resolve all spectral features, recognizing that atmospheric pressure broadening results in nearly all ro-vibrational lines having linewidths = 0.1 cm-1. As an example case, application of the database is demonstrated via identification and quantification of dominant 1H-perfluoroheptane and perfluorohept-1-ene fluorocarbon products resulting from thermal decomposition of perfluorooctanoate (PFOA) below 450 °C.

Infrared, Gas-phase spectra, FTIR, Spectral databa↗

G2Aero Database of Airfoils - Curated Airfoils

This dataset contains a curated set of 19,164 airfoil shapes from various applications and the data-driven design space of separable shape tensors (PGA space), which can be used as a parameter space for machine-learning applications focused on airfoil shapes. We constructed the airfoil dataset in two main stages. First, we identified 13 baseline airfoils from the NREL 5MW and IEA 15MW reference wind turbines. We reparameterized these shapes using least-squares fits of 8-order CST parametrizations, which involve 18 coefficients. By uniformly perturbing all 18 CST coefficients by +/-20% around each baseline airfoil, we generated 1,000 unique airfoils. Each airfoil was sampled with 1,001 shape landmarks whose x-coordinates followed a cosine distribution along the chord. This process resulted in a total of 13,000 airfoil shapes, each with 1,001 landmarks. In the second phase, we gathered additional airfoils from the extensive BigFoil database, which consolidates data from sources such as the University of Illinois Urbana-Champaign (UIUC) airfoil database, the JavaFoil database, the NACA-TR-824 database, and others. We undertook a thorough pre-processing step to filter out shapes with sparse, noisy, or incomplete data. We also removed airfoils with sharp leading edge and those exceeding our threshold for trailing edge thickness. Additionally, we thinned out the collection of NACA airfoils-- parametric sweeps of NACA airfoils with increasing thickness and camber present in BigFoil database-- by selecting every fourth step in the parameter sweeps. Finally, we regularized the airfoils by reparametrizing them with an 8-order CST parametrization (with 1,001 shape landmarks with x coordinated following cosine distribution along the chord) and removing airfoils with high reconstruction errors. This data pre-processing resulted in a set of 6,164 airfoils. In total, our curated airfoil dataset comprises 19,164 airfoils, each with 1,001 landmarks, and is stored in the curated_airfoils.npz file. Using this curated airfoil dataset, we utilized the separable shape tensors framework to develop a data-driven parameterization of airfoils based on principal geodesic analysis (PGA) of separable shape tensors. This PGA space is provided in PGAspace.npz file.

airfoils↗

A restructured and updated global soil respiration database (SRDB-V5)

Field-measured soil respiration (R S , the soil-to-atmosphere CO 2 flux) observations were compiled into a global soil respiration database (SRDB) a decade ago, a resource that has been widely used by the biogeochemistry community to advance our understanding of R S dynamics. Novel carbon cycle science questions require updated and augmented global information with better interoperability among datasets. Here, we restructured and updated the global R S database to version SRDB-V5. The updated version has all previous fields revised for consistency and simplicity, and it has several new fields to include ancillary information (e.g., R S measurement time, collar insertion depth, collar area). The new SRDB-V5 includes published papers through 2017 (800 independent studies), where total observations increased from 6633 in SRDB-V4 to 10 366 in SRDB-V5. The SRDB-V5 features more R S data published in the Russian and Chinese scientific literature and has an improved global spatio-temporal coverage and improved global climate space representation. We also restructured the database so that it has stronger interoperability with other datasets related to carbon cycle science. For instance, linking SRDB-V5 with an hourly timescale global soil respiration database (HGRsD) and a community database for continuous soil respiration (COSORE) enables researchers to explore new questions. The updated SRDB-V5 aims to be a data framework for the scientific community to share seasonal to annual field R S measurements, and it provides opportunities for the biogeochemistry community to better understand the spatial and temporal variability in R S , its components, and the overall carbon cycle.

54 ENVIRONMENTAL SCIENCES↗

Small Spacecraft Systems Virtual Institute's Federated Databases and State of the Art of Small Spacecraft Technology Report

NASA's Small Spacecraft Systems Virtual Institute (S3VI) is collaborating with the Air Force Research Laboratory and Space Dynamics Laboratory on the development of a small spacecraft parts database called SmallSat Parts On Orbit Now (SPOON). The SPOON database contains small spacecraft parts and technologies categorized by major satellite subsystems developed by industry, academia and government. The State of the Art of Small Spacecraft Technology report reflects small spacecraft parts submitted to the SPOON database and technologies compiled from other sources that were assessed as the current state of the art in each of the major subsystems. The report, first commissioned by NASA's Small Spacecraft Technology Program in mid-2013, is developed in response to the continuing growth in interest in using small spacecraft for many types of missions in Earth orbit and beyond. Due to the high market penetration of CubeSats, particular emphasis is placed on the state of the art of CubeSat-related technology. The 2018 report is planned for release in late summer. A review of SPOON database functionality, federation of additional NASA-internal and external databases along with a common search capability, as well as an overview of the State of the Art of Small Spacecraft Technology report will be presented. The S3VI is jointly sponsored by NASA's Space Technology Mission Directorate and Science Mission Directorate.

Small spacecraft↗

Dynamic IT Security Database and Analytics for Launch Control Systems Software

During the Summer 2020 session, I worked with intern Destani S. Van Arsdalen of EGS Software. Together, we co-created a tool to aid the dynamic investigation, updated over time,of the security compliance of LCS COTS and open source software. We originally planned touse spreadsheet software for management and analysis, but through this exploratoryproject, chose to use Python and JSON after receiving feedback on our project’s current anddesired capabilities at that time.At first, the project was solely designed to help on-board new COTS software, based on aquestionnaire that could be filled out for each software package. This, combined with usingthe spreadsheet application’s web-query capabilities to fetch information from the NVD,allowed presentation and analytics cells to automatically populate as elements of themanually-filled questionnaire changed. While this system was promising, we decided tochange technologies for a few reasons. In the spreadsheet, single cells could not hold complexdata like arrays and objects. The automatic population of cells and dynamic updates made itdifficult to manage and add new features. And finally, it had limited extensibility sinceadding new software required significant understanding of how both the spreadsheet wasconstructed, and the more obscure, proprietary scripting languages packaged with it.The pivot to a standard computer science database language of JSON, aided by thescripting capabilities of Python, greatly helped to improve the project’s functionality. First,and most importantly, the script’s import and analysis of database data is easilyreproducible. Additional data analysis can be modularly added without requiringmodification of the script and is capable of routine scheduling. The revised process can besplit into three parts. First, the conversion of LCS asset and software documentation into theJSON hierarchical database format. Second, the merging of this database with the NVD,forming a new data structure, using CPEs of the CVE object as a linking element betweenthem. And third, the automatically performed analytics and analysis of the combined data,in a modular and extensible format, to produce better informed business decisions. The outputted graphs, for example, are automatically generated by the Python script inconnection with the combined database. This allows updated graphs and any analytics to be re-rendered automatically following updates to the LCS’s initial asset documentation. Afinal report can then be programmatically and easily constructed from these sources to allow fully reproducible metrics for heavily evidenced risk management decisions.

it↗

UV/Vis+ Photochemistry Database: Structure, Content and Applications

The “science-softCon UV/Vis+ Photochemistry Database” (www.photochemistry.org) is a large and comprehensive collection of EUV-VUV-UV–Vis-NIR spectral data and other photochemical information assembled from published peer-reviewed papers. The database contains photochemical data including absorption, fluorescence, photoelectron, and circular and linear dichroism spectra, as well as quantum yields and photolysis related data that are critically needed in many scientific disciplines. This manuscript gives an outline regarding the structure and content of the “science-softCon UV/Vis+ Photochemistry Database”. The accurate and reliable molecular level information provided in this database is fundamental in nature and helps in proceeding further to understand photon, electron and ion induced chemistry of molecules of interest not only in spectroscopy, astrochemistry, astrophysics, Earth and planetary sciences, environmental chemistry, plasma physics, combustion chemistry but also in applied fields such as medical diagnostics, pharmaceutical sciences, biochemistry, agriculture, and catalysis. In order to illustrate this, we illustrate the use of the UV/Vis+ Photochemistry Database in four different fields of scientific endeavor.

Photochemistry↗

Developing a Sustainable, User-Friendly Literature Database to Support the Microgravity Simulation Support Facility (MSSF) at NASA's Kennedy Space Center (KSC)

Established in 2017, the Microgravity Simulation Support Facility (MSSF) at NASA’s Kennedy Space Center is the only centralized, dedicated facility supporting ground microgravity research in the United States. The MSSF offers the research community the ability to conduct simulated microgravity research with experimental conditions functionally resembling those aboard the International Space Station (ISS) and in other flight-based experimental environments. Since its inception, the MSSF has supported numerous studies and has since collected an extensive library of relevant and pertinent literature. The goal of our research was to develop and implement a sustainable, user-friendly literature database to better house this literature at the MSSF. To achieve this, our team focused on sorting, optimizing, and analyzing preexisting literature libraries to determine a best suitable and sustainable platform for the MSSF. After establishing initial database platforms, the team worked to develop descriptive and structural metadata categories to best sort the literature, which was followed by rigorous testing and optimization of the database as it was implemented. The MSSF has now been outfitted with a reliable, accessible database that effectively houses literature and provides diverse analysis to the user. Our team is continuing to test and update our platform and parameters as we aim for the formal implementation, expansion, and evolution of our database to better sustain future research ventures at the MSSF and beyond.

Database Development↗

PALMO: An OVERFLOW Machine Learning Airfoil Performance Database

The OVERFLOW Machine Learning Airfoil Performance (PALMO) database has been created to enable robust modeling of airfoil performance in a variety of applications. The database uses OVERFLOW simulation data second-order accurate in time and fourth-order accurate in space with Spalart-Allmaras turbulence closure. The foundation of the in-development PALMO database is the airfoil base cube. Each base cube includes simulation data parametrized over a range of Mach numbers, Reynolds numbers, and angles-of-attack. This first release of the database includes the NACA 4-series airfoils, with parametrization in airfoil thickness and camber from an NACA 0006 to an NACA 4424. In total, 52,480 NACA 4-series calculations were run on the NASA High-End Compute Capability (HECC) supercomputer and the corresponding airfoil performance coefficients are embedded in the Appendix of this document for public distribution. This provides high-order-accurate simulation data covering a wide range of aerospace design applications, which enables users to develop OVERFLOW-quality airfoil performance look-up tables without additional high-performance computing. In addition to engineering design and analysis of aerospace vehicles, PALMO is well suited to be a benchmark dataset for the development and testing of machine learning methods in aerospace engineering. Downstream surrogate models enable OVERFLOW- quality airfoil performance predictions for any arbitrary combination of camber, thickness, Mach number, Reynolds number, and angle-of-attack within the bounds of the database.

Database↗

A database of refractive indices and dielectric constants auto-generated using ChemDataExtractor

The ability to auto-generate databases of optical properties holds great potential for advancing optical research, especially with regards to the data-driven discovery of optical materials. An optical property database of refractive indices and dielectric constants is presented, which comprises a total of 49,076 refractive index and 60,804 dielectric constant data records on 11,054 unique chemicals. The database was auto-generated using the state-of-the-art natural language processing software, ChemDataExtractor, using a corpus of 388,461 scientific papers. The data repository offers a representative overview of the information on linear optical properties that resides in scientific papers from the past 30 years. Public availability of these data will enable a quick search for the optical property of certain materials. The large size of this repository will accelerate data-driven research on the design and prediction of optical materials and their properties. To the best of our knowledge, this is the first auto-generated database of optical properties from a large number of scientific papers. We provide a web interface to aid the use of our database.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Perovskite- and Dye-Sensitized Solar-Cell Device Databases Auto-generated Using ChemDataExtractor

The number of scientific publications reporting cutting-edge third-generation photovoltaic devices is increasing rapidly, owing to the pressing need to develop renewable-energy technologies that address the climate-change crisis. Consequently, the field could benefit from a central repository where photovoltaic-performance metrics, such as the power-conversion efficiency (η) are recorded. We present two automatically generated databases that contain photovoltaic properties and device material data for dye-sensitized solar cells (DSCs) and perovskite solar cells (PSCs), totalling 660,881 data entries representing 57,678 photovoltaic devices. The databases were generated by applying the text-mining toolkit ChemDataExtractor on a corpus of 25,720 articles. A multi-faceted evaluation, incorporating manual and automatic methods, was applied to ensure that the data contained therein were of the highest quality, with precision metrics ranging from 73.1% to 95.8%. The DSC database contains 475,045 entries representing 41,680 devices, and the PSC database contains 185,836 entries representing 15,818 devices. The databases are available in MongoDB and JSON formats, which can be queried in Python, R, Java and MATLAB for data-driven photovoltaic materials discovery.

14 SOLAR ENERGY↗

DUNE Database Development

The DUNE experiment will produce vast amounts of metadata, which describe the data coming from the read-out of the primary DUNE detectors. Various databases will make up the overall DB architecture for this metadata. ProtoDUNE at CERN is the largest existing prototype for DUNE and serves as a testing ground for - among other things - possible database solutions for DUNE. The subset of all metadata that is accessed during offline data reconstruction and analysis is referred to as ‘conditions data’ and it is stored in a dedicated database. As offline data reconstruction and analysis will be deployed on HTC and HPC resources, conditions data is expected to be accessed at very high rates. It is therefore crucial to store it in a granularity that matches the expected access patterns allowing for extensive caching. This requires a good understanding of the sources and use cases of conditions data. This contribution will briefly summarize the database architecture deployed at ProtoDUNE and explain the various sources of conditions data. We will present how the conditions data is retrieved and streamed from the databases and how it is handled to match expected access patterns.

Vizcaya Hernandez, Ana Paula↗

RNAcentral 2021: secondary structure integration, improved sequence search and new member databases

RNAcentral is a comprehensive database of non-coding RNA (ncRNA) sequences that provides a single access point to 44 RNA resources and >18 million ncRNA sequences from a wide range of organisms and RNA types. RNAcentral now also includes secondary (2D) structure information for >13 million sequences, making RNAcentral the world’s largest RNA 2D structure database. The 2D diagrams are displayed using R2DT, a new 2D structure visualization method that uses consistent, reproducible and recognizable layouts for related RNAs. The sequence similarity search has been updated with a faster interface featuring facets for filtering search results by RNA type, organism, source database or any keyword. This sequence search tool is available as a reusable web component, and has been integrated into several RNAcentral member databases, including Rfam, miRBase and snoDB. To allow for a more fine-grained assignment of RNA types and subtypes, all RNAcentral sequences have been annotated with Sequence Ontology terms. The RNAcentral database continues to grow and provide a central data resource for the RNA community. RNAcentral is freely available at https://rnacentral.org.

59 BASIC BIOLOGICAL SCIENCES↗

Mapping Inquiry Tool (MapIT) Database

The Mapping Inquiry Tool (MapIT) database consists of a geodatabase and data catalog of geologic, geophysical, structural, hydrologic, and contextual data, based on the data types to support geologic carbon storage activities and other subsurface energy systems resource assessments. The database was aggregated from publicly available data across the USA from state and federal entities. The database is structured by categories including rock unit geology, boundaries, national CS datasets, geophysical data, faults and structural data, infrastructure, surface hydrology, groundwater, and more. The data described in the data catalog is also available in the Mapping Inquiry Tool (https://edx.netl.doe.gov/dataset/mapping-inquiry-tool). Version 3 of the geodatabase and data catalog have been updated as of 5/17/2024. The database was published with a limited number of layers. The Catalog V3 contains many more resources than the geodatabase, documenting all layers that will be included in MapIT, and includes links to the original sources of the data. Within the catalog, in the final column, there is information about if the file is included in the geodatabase or not. Use the links provided in the catalog to download data directly from the original source if not included in the geodatabase. Four resources are included in this submission: 1. Geodatabase 2. ReadMe file 3. Catalog of data layers and additional data resources 4. Web link to a resource describing the motivation and reviewing the content of the geodatabase - DOE NETL Carbon Storage Site Mapping Inquiry Tool Database

carbon storage↗