Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model queries”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods.

knowledge↗

A Global Building Occupant Behavior Database

This paper introduces a database of 34 field-measured building occupant behavior datasets collected from 15 countries and 39 institutions across 10 climatic zones covering various building types in both commercial and residential sectors. This is a comprehensive global database about building occupant behavior. The database includes occupancy patterns (i.e., presence and people count) and occupant behaviors (i.e., interactions with devices, equipment, and technical systems in buildings). Brick schema models were developed to represent sensor and room metadata information. The database is publicly available, and a website was created for the public to access, query, and download specific datasets or the whole database interactively. The database can help to advance the knowledge and understanding of realistic occupancy patterns and human-building interactions with building systems (e.g., light switching, set-point changes on thermostats, fans on/off, etc.) and envelopes (e.g., window opening/closing). With these more realistic inputs of occupants’ schedules and their interactions with buildings and systems, building designers, energy modelers, and consultants can improve the accuracy of building energy simulation and building load forecasting.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Item Retrieval as Utility Estimation

Retrieval systems have greatly improved over the last half century, estimating relevance to a latent user need in a wide variety of areas. One area that has not enjoyed such advancements is searching for items by attribute values, a common activity in e-commerce and science, particularly given numeric values. Existing item retrieval systems assume the user has a firm grasp of their own desires and can formulate a good Boolean or SQL-style query to retrieve items, as one would do with a database. A contrasting approach would be to estimate how well items match the user?s latent desires and return items ranked by this estimation. Towards this end, we present a retrieval model inspired by multi-criteria decision making theory, concentrating on numeric attributes. We evaluate our novel approach, the de-facto standard of Boolean retrieval, and several models proposed in the literature, in two user studies using Amazon Mechanical Turk. We use a competitive game to motivate test subjects and compare methods based on the results of the subjects? initial query and their success in the game. In our experiments, our new method signi cantly outperformed the others, whereas the Boolean approaches had the worst performance.

operations research↗

Trustworthiness and Trust: Identifying Factors that Drive Successful Human-AI Interaction in Nuclear Power Plant Applications

Emerging technologies such as artificial intelligence (AI) and machine learning (ML) are rapidly evolving and considered a promising tool for efficient and continued safe operations of the U.S. nuclear power plants (NPPs). Emerging AI techniques like large language models (LLMs) are one such technology that may support personnel at existing NPPs perform work more efficiently. For example, operators may query the current operational status of a power plant via a chat interface leveraging LLMs to access plant-related information in an interactive manner rather than manually collecting various sensor data for tasks such as surveillances or completing work orders. This is a fundamental shift in the way operators currently perform their tasks today. The literature of human-automation interaction indicates that trust is a crucial factor that drives successful interaction between a human operator and an automated system, like an AI-infused NPP application. This work presents the results of a literature review on key factors that relate to trust in AI/LLM technologies for NPP applications. The relevant literature of human factors and cognitive engineering has identified various factors related to trust including trustworthiness, performance characteristics, operator skill and perceived risk. This preliminary literature review will guide development and evaluation of models involving the identified factors influencing trust in AI and develop a framework for human-centered design for interface between humans and AI. By addressing trust, this work supports developing a technical basis for designing key characteristics of AI/LLM to support calibrated trust, which will ultimately support wide-scale adoption of AI/LLM technologies, as well as ensure safe, effective, and reliable use.

99 - GENERAL AND MISCELLANEOUS↗

BuildStockQuery [SWR-23-58]

BuildStockQuery is a python library designed to simplify and streamline the process of querying massive, terabyte-scale datasets generated by ResStock(TM). ResStock (SWR-19-15) is a U.S. DOE-supported, NREL-built, national residential building energy stock model that enables a new approach to large-scale residential energy analysis across the U.S. by combining large public and private data sources, statistical sampling, detailed sub-hourly building simulations, and high-performance computing. BuildStockQuery offers an intuitive Object-Oriented Programming (OOP) interface to the ResStock output dataset allowing users to easily perform common queries and receive results in familiar pandas DataFrame format, abstracting away the need for complex SQL query. By initializing a query object with the pertinent Athena database and table names, users can easily query for various kinds of insights, for example, timeseries electricity for an end use for a given state grouped by building types.

Adhikari, Rajendra↗

Mapping analysis and planning system for the John F. Kennedy Space Center

Environmental management, impact assessment, research and monitoring are multidisciplinary activities which are ideally suited to incorporate a multi-media approach to environmental problem solving. Geographic information systems (GIS), simulation models, neural networks and expert-system software are some of the advancing technologies being used for data management, query, analysis and display. At the 140,000 acre John F. Kennedy Space Center, the Advanced Software Technology group has been supporting development and implementation of a program that integrates these and other rapidly evolving hardware and software capabilities into a comprehensive Mapping, Analysis and Planning System (MAPS) based in a workstation/local are network environment. An expert-system shell is being developed to link the various databases to guide users through the numerous stages of a facility siting and environmental assessment. The expert-system shell approach is appealing for its ease of data access by management-level decision makers while maintaining the involvement of the data specialists. This, as well as increased efficiency and accuracy in data analysis and report preparation, can benefit any organization involved in natural resources management.

Hall, C. R.↗

JHTDB-wind: a web-accessible large-eddy simulation database of a wind farm with virtual sensor querying

This paper introduces JHTDB-wind (https://turbulence.idies.jhu.edu/datasets/windfarms, last access: 11 November 2025), a publicly accessible database containing large-eddy simulation (LES) data from wind farms. Building on the framework of the Johns Hopkins Turbulence Database (JHTDB), which hosts direct numerical simulation (DNS) and some LES datasets of canonical turbulent flows, JHTDB-wind stores the 4D space–time history of the flow and provides users the ability to access and query the data via a web-based virtual sensor interface. The initial dataset comprises LES results from a large wind farm with 10×6 turbines, modeled using a filtered actuator line method, under conventionally neutral atmospheric conditions. These data comprise 1 h (hour) of flow field data (velocity, pressure, potential temperature deviation, subgrid-scale (SGS) eddy viscosity, and turbine forces, approximately 15 TB (terabytes) and wind turbine data – including both turbine-level operational quantities and blade-level aerodynamic quantities (approximately 1.3 TB) – stored in Zarr and Parquet formats, respectively. Data retrieval is facilitated by the giverny Python package, allowing remote users to query the database in Python or MATLAB (C and Fortran support are available for flow field data). This paper details the simulation setup and demonstrates data access through examples that analyze wind farm flow structures and turbine performance. The framework is extensible to future datasets, including the JHTDB-wind diurnal cycle simulation analyzed in Xiao et al. (2025).

17 WIND ENERGY↗

A Lightweight I/O Scheme to Facilitate Spatial and Temporal Queries of Scientific Data Analytics

In the era of petascale computing, more scientific applications are being deployed on leadership scale computing platforms to enhance the scientific productivity. Many I/O techniques have been designed to address the growing I/O bottleneck on large-scale systems by handling massive scientific data in a holistic manner. While such techniques have been leveraged in a wide range of applications, they have not been shown as adequate for many mission critical applications, particularly in data post-processing stage. One of the examples is that some scientific applications generate datasets composed of a vast amount of small data elements that are organized along many spatial and temporal dimensions but require sophisticated data analytics on one or more dimensions. Including such dimensional knowledge into data organization can be beneficial to the efficiency of data post-processing, which is often missing from exiting I/O techniques. In this study, we propose a novel I/O scheme named STAR (Spatial and Temporal AggRegation) to enable high performance data queries for scientific analytics. STAR is able to dive into the massive data, identify the spatial and temporal relationships among data variables, and accordingly organize them into an optimized multi-dimensional data structure before storing to the storage. This technique not only facilitates the common access patterns of data analytics, but also further reduces the application turnaround time. In particular, STAR is able to enable efficient data queries along the time dimension, a practice common in scientific analytics but not yet supported by existing I/O techniques. In our case study with a critical climate modeling application GEOS-5, the experimental results on Jaguar supercomputer demonstrate an improvement up to 73 times for the read performance compared to the original I/O method.

Temporal Queries↗

Latent Space Dynamics Identification

LaSDI is a data-driven physical simulation software that forms a latent space for a given high-fidelity model and discovers a set of ordinary differential equations for the latent space dynamics. It allows a fast and accurate solution process, which is useful for multi-query decision making applications, such as design optimization and uncertainty quantification. The performance of the LaSDI framework is demonstrated on four different problems, i.e., 1D and 2D Burgers equations, nonlinear heat conduction, and radial advection problems. Both linear and nonlinear compression techniques, such as neural network and proper orthogonal decomposition, are used to form a latent space. A concept of local dynamics identification procedure is introduced to enable a parametric model, which enhances the accuracy level over a given parameter space.

Fries, William↗

Evaluation of the Self Retrieval Augmented Generation Technique on Common Security Advisory Framework Data

This small experimental report evaluates a variation of Retrieval Augmented Generation (RAG), called Self-RAG. This method uses a generative language model that incorporates retrieved facts into its generation and is explicitly trained to be able to determine whether retrieved information is enough to answer the input query, with a user-defined threshold for confidence. We performed an experiment using data from the publicly available CISA Common Security Advisory Framework (CSAF) repository (https://github.com/cisagov/CSAF) as the database of facts to be used in retrieval. Qualitative results from the experiment demonstrate that the Self-RAG method has some ability to provide reasonable answers to queries that are in the dataset and will often ignore irrelevant information when asked outside of domain questions (e.g., general facts). In settings with deliberately confusing questions (the question is within domain, but asks about a fabricated advisory), it was able to refuse 40% of the time without further adjustments to the original framework. While this performance is not sufficient for current practical use, further improvements to data formatting, disambiguating results, and leveraging threshold values could improve performance significantly. However, evaluating this will require more extensive evaluations on larger datasets and potentially better models.

97 MATHEMATICS AND COMPUTING↗

Design of Digital Twin Sensing Strategies Via Predictive Modeling and Interpretable Machine Learning

This work develops a methodology for sensor placement and dynamic sensor scheduling decisions for digital twins. The digital twin data assimilation is posed as a classification problem, and predictive models are used to train optimal classification trees that represent the map from observed data to estimated digital twin states. In addition to providing a rapid digital twin updating capability, the resulting classification trees yield an interpretable mathematical representation that can be queried to inform sensor placement and sensor scheduling decisions. The proposed approach is demonstrated for a structural digital twin of a 12 ft wingspan unmanned aerial vehicle. Offline, training data are generated by simulating scenarios using predictive reduced-order models of the vehicle in a range of structural states. Furthermore, these training data can be further augmented using experimental or other historical data. In operation, the trained classifier is applied to observational data from the physical vehicle, enabling rapid adaptation of the digital twin in response to changes in structural health. Within this context, we study the performance of the optimal tree classifiers and demonstrate how they enable explainable structural assessments from sparse sensor measurements and also inform optimal sensor placement.

47 OTHER INSTRUMENTATION↗

AI to Predict Glass Compositions Satisfying Property and Cooling Rate Criteria

This project aimed to develop a predictive, artificial intelligence/machine learning-based model to identify glass compositions satisfying specified property requirements. Such a model would provide a systematic approach for narrowing down the nearly infinite range of possible compositions for glasses and minimize unnecessary experimental trial and error. A large empirical data set for training and testing the algorithm was obtained from the SciGlass database. It contains glass compositions and corresponding property data from a wide range of literature sources. However, the currently available form of this data, recently released under an open database license, is not conducive to easy querying and use. The data structure was deciphered and a customized parsing code developed to make this data more usable for the current and future work. Neural network models were developed and trained on viscosity data from the database and demonstrated potential for improving prediction accuracy over a traditional regression model.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Component-Level Inverse Design of Transmon Qubits Using Neural Networks

Designing a superconducting qubit to realize specific Hamiltonian parameters typically requires iterating through a time and compute-intensive forward loop in which the designer chooses a layout geometry, simulates it, extracts circuit parameters such as capacitances, and refines the geometry. We study the inverse version of this task using a neural-network workflow that maps target Hamiltonian parameters directly to component-level layout parameters, which we subsequently demonstrate on a planar transmon layout. During training, we pair the inverse model with a frozen forward surrogate model and evaluate the loss in Hamiltonian space rather than in layout-parameter space. In validation against a conventional EM solver, 97% of generated designs produce usable geometries, and the inverse-plus-surrogate pipeline reaches mean percent errors of 0.73% for qubit frequency and 1.58% for anharmonicity, comparable to or below the fabrication and simulation-to-measurement uncertainty expected for academic-process transmon devices of this type. A single pipeline query takes ~60 ms on CPU, versus ~2 min for a conventional EM capacitance extraction on the same hardware, a speedup of approximately 2,000x. Batching minimizes the AI model inference overhead, reducing the runtime to 3.1 microseconds per sample on CPU and 2.6 microseconds per sample on GPU at a batch size of 2048, resulting in speedups of 3.9 x 10^7 and 4.6 x 10^7, respectively, relative to a single conventional CPU EM extraction. Our results indicate that component-level inverse design usefully extends and complements conventional EM simulation, including for small datasets on the order of 1,000 samples.

Seidel, Olivia [Fermilab; Texas U., Arlington]↗

Virtual VMASC: A 3D Game Environment

The advantages of creating interactive 3D simulations that allow viewing, exploring, and interacting with land improvements, such as buildings, in digital form are manifold and range from allowing individuals from anywhere in the world to explore those virtual land improvements online, to training military personnel in dealing with war-time environments, and to making those land improvements available in virtual worlds such as Second Life. While we haven't fully explored the true potential of such simulations, we have identified a requirement within our organization to use simulations like those to replace our front-desk personnel and allow visitors to query, naVigate, and communicate virtually with various entities within the building. We implemented the Virtual VMASC 3D simulation of the Virginia Modeling Analysis and Simulation Center (VMASC) office building to not only meet our front-desk requirement but also to evaluate the effort required in designing such a simulation and, thereby, leverage the experience we gained in future projects of this kind. This paper describes the goals we set for our implementation, the software approach taken, the modeling contribution made, and the technologies used such as XNA Game Studio, .NET framework, Autodesk software packages, and, finally, the applicability of our implementation on a variety of architectures including Xbox 360 and PC. This paper also summarizes the result of our evaluation and the lessons learned from our effort.

Manepalli, Suchitra↗

Semantic Search with Sentence-BERT for Design Information Retrieval

Managing and referencing design knowledge is a critical activity in the design process. However, reliably retrieving useful knowledge can be a frustrating experience for users of knowledge management systems due to inherent limitations of standard keyword-based searches. In this research, we consider the task of retrieving relevant lessons learned from the NASA Lessons Learned Information System (LLIS). To this end, we apply a state-of-the-art natural language processing (NLP) technique for information retrieval: semantic search with sentence-BERT, which is a modification of a Bidirectional Encoder Representations from Transformers (BERT) model that uses siamese and triplet network architectures to obtain semantically meaningful sentence embeddings. While the pretrained sBERT model shows excellent out-of-the-box performance, we further fine-tune the model on data from the LLIS so that it learns on design engineering-relevant vocabulary. We quantify the improvement in query results using both standard sBERT and fine-tuned sBERT over the LLIS’s built-in keyword search. Additionally, we demonstrate a use case for the query system by searching for lessons learned relevant to specific requirements from a NASA project as part of a broader knowledge management and retrieval system. Results indicate that applying state-of-the-art natural language processing techniques, especially when fine-tuned using engineering data, to design information retrieval tasks shows significant promise in modernizing design knowledge management systems.

Hannah S. Walsh↗

Characterizing Tradeoffs in Memory, Accuracy, and Speed for Chemistry Tabulation Techniques

Chemistry tabulation is a common approach in practical simulations of turbulent combustion at engineering scales. Linear interpolants have traditionally been used for accessing precomputed multidimensional tables but suffer from large memory requirements and discontinuous derivatives. Higher-degree interpolants address some of these restrictions but are similarly limited to relatively low-dimensional tabulation. Artificial neural networks (ANNs) can be used to overcome these limitations but cannot guarantee the same accuracy as interpolants and introduce challenges in reproducibility and reliable training. These challenges are enhanced as the physics complexity to be represented within the tabulation increases. Here, we assess the efficiency, accuracy, and memory requirements of Lagrange polynomials, tensor product B-splines, and ANNs as tabulation strategies. We analyze results in the context of nonadiabatic flamelet modeling where higher dimension counts are necessary. While ANNs do not require structuring of data, providing benefits for complex physics representation, interpolation approaches often rely on some structuring of the table. Interpolation using structured table inputs that are not directly related to the variables transported in a simulation can incur additional query costs. This is demonstrated in the present implementation of heat losses. We show that ANNs, despite being difficult to train and reproduce, can be advantageous for high-dimensional, unstructured datasets relevant to nonadiabatic flamelet models. Furthermore we demonstrate that Lagrange polynomials show significant speedup for similar accuracy compared to B-splines.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Background-Aware 3-D Point Cloud Segmentation With Dynamic Point Feature Aggregation

With the proliferation of LiDAR sensors and 3-D vision cameras, 3-D point cloud analysis has attracted significant attention in recent years. In this article, we propose a novel 3-D point cloud learning network, referred to as dynamic point feature aggregation network (DPFA-Net), by selectively performing the neighborhood feature aggregation (FA) with dynamic pooling and an attention mechanism. DPFA-Net has two variants for semantic segmentation and classification of 3-D point clouds. As the core module of the DPFA-Net, we propose an FA layer, in which features of the dynamic neighborhood of each point are aggregated via a self-attention mechanism. In contrast to other segmentation models, which aggregate features from fixed neighborhoods, our approach can aggregate features from different neighbors in different layers providing a more selective and broader view to the query points and focusing more on the relevant features in a local neighborhood. In addition, to further improve the performance of semantic segmentation, we exploit the background–foreground (BF) information and present two novel approaches, namely, two-stage BF-Net and BF regularization. Experimental results show that the proposed DPFA-Net achieves the state-of-the-art overall accuracy score of 89.22% for semantic segmentation on the Stanford large-scale 3-D Indoor Spaces (S3DIS) dataset and provides consistently satisfactory performance across different tasks of semantic segmentation, part segmentation, and 3-D object classification. Furthermore, our model achieves 93.1% accuracy on the ModelNet40 dataset and provides a mean shape intersection-over-union (IoU) value of 85.5% for part segmentation on the ShapeNet-Part dataset. It is a also computationally more efficient compared to other methods.

3-D↗

Practical galaxy morphology tools from deep supervised representation learning

Astronomers have typically set out to solve supervised machine learning problems by creating their own representations from scratch. We show that deep learning models trained to answer every Galaxy Zoo DECaLS question learn meaningful semantic representations of galaxies that are useful for new tasks on which the models were never trained. We exploit these representations to outperform several recent approaches at practical tasks crucial for investigating large galaxy samples. The first task is identifying galaxies of similar morphology to a query galaxy. Given a single galaxy assigned a free text tag by humans (e.g. ‘#diffuse’), we can find galaxies matching that tag for most tags. The second task is identifying the most interesting anomalies to a particular researcher. Our approach is 100 per cent accurate at identifying the most interesting 100 anomalies (as judged by Galaxy Zoo 2 volunteers). The third task is adapting a model to solve a new task using only a small number of newly labelled galaxies. Models fine-tuned from our representation are better able to identify ring galaxies than models fine-tuned from terrestrial images (ImageNet) or trained from scratch. We solve each task with very few new labels; either one (for the similarity search) or several hundred (for anomaly detection or fine-tuning). This challenges the longstanding view that deep supervised methods require new large labelled data sets for practical use in astronomy. To help the community benefit from our pretrained models, we release our fine-tuning code zoobot. Zoobot is accessible to researchers with no prior experience in deep learning.

79 ASTRONOMY AND ASTROPHYSICS↗