Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “document repository”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

MIKA: Manager for Intelligent Knowledge Access Toolkit for Engineering Knowledge Discovery and Information Retrieval

Repositories of safety reports are often underutilized and only analyzed manually by trained experts, despite safety management systems requiring reports. These collections of documents contain a wealth of information from past projects and operations that could improve system safety and design. Advances in natural language processing techniques have improved information extraction and retrieval in consumer technology, biomedicine, and finance, for instance, but have not been applied to engineering documents on the same scale. To this end, the Manager for Intelligent Knowledge Access (MIKA) open-source toolkit has been developed for rapid knowledge discovery and information retrieval in safety engineering applications. The MIKA toolkit uses state-of-the-art natural language processing algorithms and allows a user to apply these methods to their own dataset. This paper describes the MIKA toolkit and its two primary capabilities, knowledge discovery and information retrieval, and demonstrates the toolkit via a case study on National Transportation Safety Board (NTSB) reports.

Machine Learning↗

MIKA: Manager for Intelligent Knowledge Access Toolkit for Engineering Knowledge Discovery and Information Retrieval

Repositories of safety reports are often underutilized and only analyzed manually by trained experts, despite safety management systems requiring reports. These collections of documents contain a wealth of information from past projects and operations that could improve system safety and design. Advances in natural language processing techniques have improved information extraction and retrieval in consumer technology, biomedicine, and finance, for instance, but have not been applied to engineering documents on the same scale. To this end, the Manager for Intelligent Knowledge Access (MIKA) open-source toolkit has been developed for rapid knowledge discovery and information retrieval in safety engineering applications. The MIKA toolkit uses state-of-the-art natural language processing algorithms and allows a user to apply these methods to their own dataset. This paper describes the MIKA toolkit and its two primary capabilities, knowledge discovery and information retrieval, and demonstrates the toolkit via a case study on National Transportation Safety Board (NTSB) reports.

Systems Engineering↗

The Nasa SRA Process as It Relates to Open-Source Workflows Developed for GeneLab Data Processing

To release open, standards-compliant processed data sets in the Open Science Data Repository (OSDR), the GeneLab Data Processing team works with the scientific community through the OSDR Analysis Working Groups to design and build open-source data processing pipelines. Once baselined internally, these pipelines are wrapped into workflows and published on the NASA GeneLab Data Processing public GitHub repository along with detailed instructions for installation and use. Each workflow must be approved through NASA's Software Release Authorization (SRA) process prior to publishing. However, the SRA process lacks sufficient documentation and clarity regarding which forms are applicable for new open-source software that utilizes publicly available 3rd party tools, and the SRA process can take several months to complete, making sharing software outside of NASA cumbersome and in contradiction with the concept of Open Science. Furthermore, the SRA process was designed as a one-size fits all approach and thus many of the questions asked are not applicable to our open-source workflows. Here we describe the software provided on the NASA GeneLab Data Processing GitHub repository, summarize our experiences with the SRA process to release these software, and propose a more stream-lined approach for review of open-source projects.

Software Release Authorization↗

The NASA SRA Process as it Relates to Open-Source Workflows Developed for GeneLab Data Processing

To release open, standards-compliant processed data sets in the Open Science Data Repository (OSDR), the GeneLab Data Processing team works with the scientific community through the OSDR Analysis Working Groups to design and build open-source data processing pipelines. Once baselined internally, these pipelines are wrapped into workflows and published on the NASA GeneLab Data Processing public GitHub repository along with detailed instructions for installation and use. Each workflow must be approved through NASA's Software Release Authorization (SRA) process prior to publishing. However, the SRA process lacks sufficient documentation and clarity regarding which forms are applicable for new open-source software that utilizes publicly available 3rd party tools, and the SRA process can take several months to complete, making sharing software outside of NASA cumbersome and in contradiction with the concept of Open Science. Furthermore, the SRA process was designed as a one-size fits all approach and thus many of the questions asked are not applicable to our open-source workflows. Here we describe the software provided on the NASA GeneLab Data Processing GitHub repository, summarize our experiences with the SRA process to release these software, and propose a more stream-lined approach for review of open-source projects.

Software Release Authorization↗

ResStock™ v3.2.0 [SWR-19-15 and SWR-20-07]

The ResStock™ analysis tool was built on NREL's OpenStudio® platform, and is a project geared at modeling existing residential building stocks at national, regional, or local scales with a high-degree of granularity (e.g., one physics-based simulation model for every 200 dwelling units), using the EnergyPlus® simulation engine. Information about ComStock™, a sister tool for modeling the commercial building stock, can be found here: https://www.nrel.gov/buildings/comstock.html This repository contains: Housing characteristics of the U.S. residential building stock, in the form of conditional probability distributions stored as tab-separated value (.tsv) files. Comments at the bottom of each file document data sources and assumptions for each. A library of housing characteristic "options" that translate high-level characteristic parameters into arguments for OpenStudio measures, and which are referenced by the housing characteristic .tsv files and building energy upgrades defined in project definition files Project definition files: v2.3.0 and later: buildstockbatch YML files openable in any text editor v2.2.5 and prior: Project folder openable in PAT Unit-level OpenStudio Measures for automatically constructing OpenStudio Models of each representative dwelling unit model: v3.0.0 and later: OpenStudio-HPXML Measures v2.5.0 and prior: OpenStudio Measures Higher-level OpenStudio Measures for controlling simulation inputs and outputs This repository does not contain software for running ResStock simulations, which can be found as follows: Versions 2.3.0 and later only support the use of buildstockbatch for deploying simulations on high-performance or cloud computing. Version 2.3.0 also removed separate projects for single-family detached and multifamily buildings, in lieu of a combined project_national representing the U.S. residential building stock. See the changelog for more details. Versions 2.2.5 and prior support the use of the publicly available OpenStudio-PAT software as an interface for deploying simulations on cloud computing. Read the documentation for v2.2.5.

Horowitz, Scott↗

Resource Recovery for the Wastewater Industry

This information sheet discusses the technology pillar, Resource Recovery, as a pathway toward improving wastewater infrastructure sustainability and resiliency. To supplement existing literature on current technologies and policies for improving resiliency at wastewater (WW) treatment plants, this document aims to accomplish the following: • Summarize wastewater sludge recovery methods • Summarize biogas production and codigestion methods • Serve as a comprehensive (though not exhaustive) repository for resource recovery for wastewater utilities The Resource Recovery Technical Information Sheet should be viewed as a general guide to established best practices for the water and wastewater (W/WW) sector when considering implementing energy capture technologies. Additional details on associated energy capture avenues such as combined heat and power (CHP), renewable energy, and inline hydropower from tertiary effluent in W/WW facilities are presented in the Energy Capture Technology Information Sheet.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Bayesian Inference for NASA Probabilistic Risk and Reliability Analysis

This document, Bayesian Inference for NASA Probabilistic Risk and Reliability Analysis, is intended to provide guidelines for the collection and evaluation of risk and reliability-related data. It is aimed at scientists and engineers familiar with risk and reliability methods and provides a hands-on approach to the investigation and application of a variety of risk and reliability data assessment methods, tools, and techniques. This document provides both: A broad perspective on data analysis collection and evaluation issues. A narrow focus on the methods to implement a comprehensive information repository. The topics addressed herein cover the fundamentals of how data and information are to be used in risk and reliability analysis models and their potential role in decision making. Understanding these topics is essential to attaining a risk informed decision making environment that is being sought by NASA requirements and procedures such as 8000.4 (Agency Risk Management Procedural Requirements), NPR 8705.05 (Probabilistic Risk Assessment Procedures for NASA Programs and Projects), and the System Safety requirements of NPR 8715.3 (NASA General Safety Program Requirements).

Dezfuli, Homayoon↗

SeaWiFS technical report series. Volume 20: The SeaWiFS bio-optical archive and storage system (SeaBASS), part 1

This document provides an overview of the Sea-viewing Wide Field-of-view Sensor (SeaWiFS) Bio-Optical Archive and Storage System (SeaBASS), which will serve as a repository for numerous data sets of interest to the SeaWiFS Science Team and other approved investigators in the oceanographic community. The data collected will be those data sets suitable for the development and evaluation of bio-optical algorithms which include results from SeaWiFS Intercalibration Round-Robin Experiments (SIRREXs), prelaunch characterization of the SeaWiFS instrument by its manufacturer -- Hughes/Santa Barbara Research Center (SBRC), Marine Optical Characterization Experiment (MOCE) cruises, Marine Optical Buoy (MOBY) deployments and refurbishments, and field studies of other scientists outside of NASA. The primary goal of the data system is to provide a simple mechanism for querying the available archive and requesting specific items, while assuring that the data is made available only to authorized users. The design, construction, and maintenance of SeaBASS is the responsibility of the SeaWiFS Calibration and Validation Team (CVT). This report is concerned with documenting the execution of this task by the CVT and consists of a series of chapters detailing the various data sets involved. The topics presented are as follows: 1) overview of the SeaBASS file architecture, 2) the bio-optical data system, 3) the historical pigment database, 4) the SIRREX database, and 5) the SBRC database.

Hooker, Stanford B.↗

Phenopacket-tools: Building and validating GA4GH Phenopackets

The Global Alliance for Genomics and Health (GA4GH) is a standards-setting organization that is developing a suite of coordinated standards for genomics. The GA4GH Phenopacket Schema is a standard for sharing disease and phenotype information that characterizes an individual person or biosample. The Phenopacket Schema is flexible and can represent clinical data for any kind of human disease including rare disease, complex disease, and cancer. It also allows consortia or databases to apply additional constraints to ensure uniform data collection for specific goals. We present phenopacket-tools, an open-source Java library and command-line application for construction, conversion, and validation of phenopackets. Phenopacket-tools simplifies construction of phenopackets by providing concise builders, programmatic shortcuts, and predefined building blocks (ontology classes) for concepts such as anatomical organs, age of onset, biospecimen type, and clinical modifiers. Phenopacket-tools can be used to validate the syntax and semantics of phenopackets as well as to assess adherence to additional user-defined requirements. The documentation includes examples showing how to use the Java library and the command-line tool to create and validate phenopackets. We demonstrate how to create, convert, and validate phenopackets using the library or the command-line application. Source code, API documentation, comprehensive user guide and a tutorial can be found at https://github.com/phenopackets/phenopacket-tools. The library can be installed from the public Maven Central artifact repository and the application is available as a standalone archive. The phenopacket-tools library helps developers implement and standardize the collection and exchange of phenotypic and other clinical data for use in phenotype-driven genomic diagnostics, translational research, and precision medicine applications.

59 BASIC BIOLOGICAL SCIENCES↗

Using Parameter Sweep in WaterTAP to Analyze New Water Treatment Technologies

We describe a powerful and generalized parameter sweep tool in this report that was originally developed to analyze the performance of existing and novel water treatment models being developed in WaterTAP. Since WaterTAP is built upon IDAES and Pyomo, the parameter sweep tool can be used to systematically explore and debug the behavior of most Pyomo and IDAES numerical models. In order to enable meaningful analyses, the parameter sweep tool has been designed with the following features: 1) Model flexibility: The parameter sweep tool does not enforce any restrictions on the types of models that can be used with it. As long as a Pyomo model can be solved and the parameter is active and mutable, the tool only needs functions that describe how to run the model, the sweep parameters, and the output quantities of interest. 2) Flexible sampling: The parameter sweep tool has inbuilt functions to generate samples from a random distribution or a multidimensional Euclidean space. Furthermore, the users have to ability to supply samples generated from a tool of their choice. 3) Multiple sweep types: A user can choose from one of 3 types of parameter sweeps depending on their needs. 4) Detailed outputs: Outputs generated by the parameter sweep tool can be stored in detailed H5 file or user-friendly CSV files for post processing. 5) Parallel computing: The parameter sweep supports shared and distributed memory parallel computing to enable the use of high performance computers (HPC) for large-scale analyses. 6) Modular: The parameter sweep tool is self-contained and can easily be integrated within an outer-loop analysis or as desired by the user. 7) Ease of use: The tool is well documented and a simple sweep can be easily executed by following the online documentation in a few lines of code. We demonstrate the use of the parameter sweep tool on a simple water treatment system from the WaterTAP repository and show its parallel scaling performance on an Apple laptop and NREL's Eagle HPC. The parameter sweep tool is actively being used with models currently being developed within WaterTAP and we expect its use to grow beyond it to other IDAES and Pyomo models.

97 MATHEMATICS AND COMPUTING↗

Globally Gridded Groundwater Extraction Volumes and Costs under Six Depletion and Ponded Depth Targets

This repository contains simulated outputs from superwell – a hydro-economic tool for long-term assessment of groundwater cost and supply – providing globally gridded groundwater extractable volumes and associated unit costs ($/km³) for accessible groundwater production, based on a variety of user-defined depletion and ponded depth scenarios. Key model documentation: Niazi, H., Ferencz, S. B., Graham, N. T., Yoon, J., Wild, T. B., Hejazi, M., Watson, D. J., & Vernon, C. R. (2025). Long-term hydro-economic analysis tool for evaluating global groundwater cost and supply: Superwell v1.1. Geoscientific Model Development, 18(5), 1737-1767. https://doi.org/10.5194/gmd-18-1737-2025 Find the source code of the superwell model on GitHub: https://github.com/JGCRI/superwell Repository Overview Main output: superwell_outputs.7z contains 6 files (4.5 GB) named as superwell_py_deep_all_0.*PD_0.*DL.csv. These files present superwell outputs of global groundwater extraction volumes and cost estimates on a 0.5° scale for six scenarios with different Ponded Depth (PD; 0.3 and 0.6 m) and Depletion Limit (DL; 5%, 25%, and 40% of available volume) targets over the entire pumping lifetime of a grid cell superwell_py_deep_all_0.3PD_0.25DL_sample_100.csv contains superwell outputs for 100 data points sampled to match the global inputs' distribution superwell_py_deep_all_0.3PD_0.25DL_Grid_72548.csv contains superwell output for a single grid cell concept_v5.png provides an overview of the superwell workflow Outputs Description year_number: year of pumping depletion_limit: set depletion limit (DL) as a volume fraction of total available groundwater Mappings: continent, country, gcam_basin_id, Basin_long_name, grid_id: geographic identifiers and basin information Inputs: grid_area (km²): area of the grid cell whyclass: hydrogeological classification of the aquifer permeability (m/day), porosity (%), total_thickness (m), depth_to_water (m): aquifer properties. The geo-processed input data has been published separately: https://doi.org/10.57931/2307831 Model outputs: orig_aqfr_sat_thickness (m), aqfr_sat_thickness (m): original and remaining/instantaneous saturated thickness of the aquifer hydraulic_conductivity (m/day), transmissivity (m²/day): hydraulic properties of the aquifer radius_of_influence (m), areal_extent (km²): well radius and area of influence from the center of the well number_of_wells (-): number of wells in a grid cell determined by a ratio of well area and grid area max_drawdown (m), drawdown (m), drawdown_interference (m): well and aquifer drawdown during extraction total_head (m): total lift for the groundwater (depth to water plus drawdown) total_well_length (m): total depth of wells drilled well_yield (m³/day): pumping rate or well yield power (kW), energy (kWh): power and energy required for pumping groundwater Volume Outputs: volume_produced_perwell (m³), cumulative_vol_produced_perwell (m³): production volume metrics per well volume_produced_allwells (m³), cumulative_vol_produced_allwells (m³): aggregate extraction volumes for all wells in a grid cell available_volume (m³): available groundwater in storage for the grid cell as determined by aquifer properties depleted_vol_fraction: fraction of total volume pumped over available volumes in a grid cell (same as depletion limit) Cost Outputs: well_installation_cost ($): well installation cost based on the hydrogeological complexity of the aquifer annual_capital_cost, maintenance_cost, nonenergy_cost ($): nonenergy costs energy_cost_rate ($/kWh): electricity rate energy_cost ($): energy cost of pumping groundwater total_cost_perwell ($), total_cost_allwells ($): total annual energy and non-energy cost for each and all wells in a grid cell a unit_cost ($/m³), unit_cost_per_km3 ($/km³), unit_cost_per_acreft ($/acre-ft): total cost of pumping a unit of groundwater, indicated for different spatial units Key Resources Model documentation: Niazi, H., Ferencz, S., Graham, N., Yoon, J., Wild, T., Hejazi, M., Watson, D., & Vernon, C. (2024; In-prep). Long-term Hydro-economic Assessment Tool for Evaluating Global Groundwater Cost and Supply: Superwell v1. Geoscientific Model Development. Input data: Niazi, H., Watson, D., Hejazi, M., Yonkofski, C., Ferencz, S., Vernon, C., Graham, N., Wild, T., & Yoon, J. (2024). Global Geo-processed Data of Aquifer Properties by 0.5° Grid, Country and Water Basins. MSD-LIVE Data repository. https://doi.org/10.57931/2307831 superwell source code: https://github.com/JGCRI/superwell Cite as Niazi, H., Ferencz, S., Yoon, J., Graham, N., Wild, T., Hejazi, M., Watson, D., & Vernon, C. (2024). Globally Gridded Groundwater Extraction Volumes and Costs under Six Depletion and Ponded Depth Targets. MSD-LIVE Data repository. https://doi.org/10.57931/2307832 Contact Reach out to Hassan Niazi or Stephen Ferencz or open an issue in the superwell repository for questions or suggestions.

Earth Systems↗

ISHM Anomaly Lexicon for Rocket Test

Integrated Systems Health Management (ISHM) is a comprehensive capability. An ISHM system must detect anomalies, identify causes of such anomalies, predict future anomalies, help identify consequences of anomalies for example, suggested mitigation steps. The system should also provide users with appropriate navigation tools to facilitate the flow of information into and out of the ISHM system. Central to the ability of the ISHM to detect anomalies is a clearly defined catalog of anomalies. Further, this lexicon of anomalies must be organized in ways that make it accessible to a suite of tools used to manage the data, information and knowledge (DIaK) associated with a system. In particular, it is critical to ensure that there is optimal mapping between target anomalies and the algorithms associated with their detection. During the early development of our ISHM architecture and approach, it became clear that a lexicon of anomalies would be important to the development of critical anomaly detection algorithms. In our work in the rocket engine test environment at John C. Stennis Space Center, we have access to a repository of discrepancy reports (DRs) that are generated in response to squawks identified during post-test data analysis. The DR is the tool used to document anomalies and the methods used to resolve the issue. These DRs have been generated for many different tests and for all test stands. The result is that they represent a comprehensive summary of the anomalies associated with rocket engine testing. Fig. 1 illustrates some of the data that can be extracted from a DR. Such information includes affected transducer channels, narrative description of the observed anomaly, and the steps used to correct the problem. The primary goal of the anomaly lexicon development efforts we have undertaken is to create a lexicon that could be used in support of an associated health assessment database system (HADS) co-development effort. There are a number of significant byproducts of the anomaly lexicon compilation effort. For example, (1) Allows determination of the frequency distribution of anomalies to help identify those with the potential for high return on investment if included in automated detection as part of an ISHM system, (2) Availability of a regular lexicon could provide the base anomaly name choices to help maintain consistency in the DR collection process, and (3) Although developed for the rocket engine test environment, most of the anomalies are not specific to rocket testing, and thus can be reused in other applications.

Schmalzel, John L.↗

Solubility and Dissolution Rate of LiCl-KCl-NaCl

This work aims to determine the potential risk of directly storing waste salt from the electrorefining process of used nuclear fuel in a geologic repository. To accomplish this, the solubility limit and dissolution rate of four representative chloride salt mixtures (solutes) in water and two brine solutions (solvents) are observed and documented.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

SEED: Semantic Energy Exploration and Discovery

The Bioenergy Knowledge Discovery Framework (KDF) hosts a vast repository of specialized data, yet traditional keyword-based search methods often struggle to provide direct answers, requiring significant domain expertise and manual effort to filter through raw documents. To overcome these barriers, this software introduces a semantic search engine that enables both specialists and non-specialists to query the KDF using natural language. By shifting from rigid keyword matching to intent-based retrieval, the tool automatically identifies and ranks the most relevant sources within the database. The system functions by processing natural language queries to extract the most pertinent information, delivering an AI-generated plain-language summary alongside exact supporting quotes from retrieved documents. This integrated approach provides users with immediate, evidence-based answers while eliminating the need for exhaustive manual review. By surfacing direct insights and contextual evidence, the software enhances the usability of existing KDF resources and democratizes access to complex bioenergy data. Ultimately, this semantic search solution accelerates the discovery process and supports faster, more informed decision-making across the bioenergy sector.

Pan, Meiyu (Melrose) [Oak Ridge National Laborator↗

NSSDC and WDC-A-R/S document availability and distribution services

Documents available from the National Space Science Data Center (NSSDC) and the World Data Center A for Rockets Satellites are described. The availability, costs, ordering procedures for documents presently available, and the procedures for obtaining future documents are given. NSSDC, established by NASA to further the widest practicable use of reduced data obtained from space science investigations and to provide investigators with an active repository for such data, is responsible for the active collection, organization, storage, announcement, retrieval, dissemination, and exchange of data received from satellite experiments. Information on sounding rocket investigations is also collected.

Source record↗

Recent technology products from Space Human Factors research

The goals of the NASA Space Human Factors program and the research carried out concerning human factors are discussed with emphasis given to the development of human performance models, data, and tools. The major products from this program are described, which include the Laser Anthropometric Mapping System; a model of the human body for evaluating the kinematics and dynamics of human motion and strength in microgravity environment; an operational experience data base for verifying and validating the data repository of manned space flights; the Operational Experience Database Taxonomy; and a human-computer interaction laboratory whose products are the display softaware and requirements and the guideline documents and standards for applications on human-computer interaction. Special attention is given to the 'Convoltron', a prototype version of a signal processor for synthesizing the head-related transfer functions.

Jenkins, James P.↗

Data-Driven Buy Clean: Decarbonization and Beyond

This report was compiled to provide recommendations on the availability of public background data from the U.S. Federal life cycle assessment (LCA) Data Commons to be conformant with the Association for Life Cycle Assessment (ACLCA) 2022 Product Category Rule (PCR) Open Standard to build technical tools that can assist industry in creating more comparable Type II Environmental Product Declarations (EPDs) for Federal Buy Clean and sustainability initiatives. The Federal LCA Commons is not only a public data source but also a consistently structured, self-referencing mega-repository for data developed by federal agency experts (in agency repositories) and by academia, nonprofit organizations, and industry (via the US Life Cycle Inventory Database). The Federal LCA Commons Technical Working Group is continuously improving the standardization of data documentation, formatting, and nomenclature to ensure lossless data loading and accurate data representation. This report and appendixes include the following: 1) An introduction to data-driven Buy Clean and decarbonization initiatives at the federal level; 2) The current status and associated challenges with LCA data and EPD standards and comparability; 3) Opportunities for the Federal LCA Commons to support conformance with the ACLCA 2022 PCR Open Standard and provide resources to implement the Federal Sustainability Plan, Buy Clean Program, and Inflation Reduction Act (IRA) sustainability goals and objectives. To date, the Federal LCA Commons is the result of coordinated work by National Renewable Energy Laboratory (NREL), the U.S. Department of Agriculture (USDA), the Environmental Protection Agency (EPA), the National Energy Technology Laboratory (NETL), the Argonne National Laboratory (ANL), the U.S. Army Corps of Engineers (USACE), the Federal Highway Administration (FHWA), the U.S. Forest Service (USFS), the Federal Aviation Administration (FAA), the Department of Defense (DoD) and the National Institute of Standards and Technologies (NIST). The Federal LCA Commons will continue to combine databases from the collaborating agencies while remaining a public resource. There are several initiatives among the collaborating agencies to expand the Federal LCA Commons and dedicated federal funding and resources could accelerate and strengthen these initiatives.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗