Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model queries”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Radionuclide-specific Parameters Dataset

The radionuclide-specific parameters dataset is searchable for radiological information for multiple isotopes simultaneously. After selecting radionuclides of interest and the desired parameters, the RAIS will generate a table containing the values, chosen according to an established hierarchy. Results can be downloaded in Excel format. 50 parameters are available, including atomic number, soil to animal transfer coefficients, plant uptake coefficients, half-life, specific activity, and water solubility. Seven primary sources are used to populate the dataset of radiological-specific parameters. These values should be used in cancer risk assessments for the calculation of preliminary remediation goals (PRGs), hazard characterization, and transport modeling. Users can select up to 1000 radionuclides per query. The dataset supports environmental risk assessments, regulatory decision-making, and environmental planning with tools for benchmarking against risk-based standards. This structured approach ensures a robust evaluation of environmental risks tailored to regulatory needs.

Manning, Karessa [Oak Ridge National Laboratory (O↗

High-Dimensional Bayesian Optimization via Semi-Supervised Learning with Optimized Unlabeled Data Sampling

We introduce a novel semi-supervised learning approach, named Teacher-Student Bayesian Optimization (TSBO ), integrating the teacher-student paradigm into BO to minimize expensive labeled data queries for the first time. TSBO incorporates a teacher model, an unlabeled data sampler, and a student model. The student is trained on unlabeled data locations generated by the sampler, with pseudo labels predicted by the teacher. The interplay between these three components implements a unique selective regularization to the teacher in the form of student feedback. This scheme enables the teacher to predict high-quality pseudo labels, enhancing the generalization of the GP surrogate model in the search space. To fully exploit TSBO , we propose two optimized unlabeled data samplers to construct effective student feedback that well aligns with the objective of Bayesian optimization. Furthermore, we quantify and leverage the uncertainty of the teacher-student model for the provision of reliable feedback to the teacher in the presence of risky pseudo-label predictions. TSBO demonstrates significantly improved sample-efficiency in several global optimization tasks under tight labeled data budgets. The implementation is available at https://github.com/reminiscenty/TSBO-Official.

Yin, Yuxuan↗

Profiling the BLAST bioinformatics application for load balancing on high-performance computing clusters

Abstract Background The Basic Local Alignment Search Tool (BLAST) is a suite of commonly used algorithms for identifying matches between biological sequences. The user supplies a database file and query file of sequences for BLAST to find identical sequences between the two. The typical millions of database and query sequences make BLAST computationally challenging but also well suited for parallelization on high-performance computing clusters. The efficacy of parallelization depends on the data partitioning, where the optimal data partitioning relies on an accurate performance model. In previous studies, a BLAST job was sped up by 27 times by partitioning the database and query among thousands of processor nodes. However, the optimality of the partitioning method was not studied. Unlike BLAST performance models proposed in the literature that usually have problem size and hardware configuration as the only variables, the execution time of a BLAST job is a function of database size, query size, and hardware capability. In this work, the nucleotide BLAST application BLASTN was profiled using three methods: shell-level profiling with the Unix “time” command, code-level profiling with the built-in “profiler” module, and system-level profiling with the Unix “gprof” program. The runtimes were measured for six node types, using six different database files and 15 query files, on a heterogeneous HPC cluster with 500+ nodes. The empirical measurement data were fitted with quadratic functions to develop performance models that were used to guide the data parallelization for BLASTN jobs. Results Profiling results showed that BLASTN contains more than 34,500 different functions, but a single function, RunMTBySplitDB, takes 99.12% of the total runtime. Among its 53 child functions, five core functions were identified to make up 92.12% of the overall BLASTN runtime. Based on the performance models, static load balancing algorithms can be applied to the BLASTN input data to minimize the runtime of the longest job on an HPC cluster. Four test cases being run on homogeneous and heterogeneous clusters were tested. Experiment results showed that the runtime can be reduced by 81% on a homogeneous cluster and by 20% on a heterogeneous cluster by re-distributing the workload. Discussion Optimal data partitioning can improve BLASTN’s overall runtime 5.4-fold in comparison with dividing the database and query into the same number of fragments. The proposed methodology can be used in the other applications in the BLAST+ suite or any other application as long as source code is available.

59 BASIC BIOLOGICAL SCIENCES↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

An Open-source Llm Enhanced-tool Specialized In Helping Moose Related Problems And Tasks

MOOSEenger is an open-source, terminal-first chat application for the MOOSE ecosystem that couples specialized parsing of MOOSE documentation and “.i” input files with retrieval-augmented generation to deliver grounded answers about multiphysics modeling and workflows. It includes dedicated readers for MOOSE-style HTML and a pyhit-based parser that uses the MOOSE syntax tree to preserve block structure and attach retrieval metadata. A data-ingestion pipeline performs semantic chunking into atomic facts and stores them hierarchically in a local Chroma vector database that maintains parent–child relationships across documents; the system can ingest directories, individual files, and single-page web content, and it provides CRUD operations (insert, update, delete) to manage the corpus. At query time, relevant chunks are embedded, retrieved, and fused into the model context, with interactive features such as token streaming, persistent chat history, and dynamic RAG (retrieval triggered by user input or intermediate model output). Deployment is flexible: MOOSEenger runs with local Ollama models or remote Hugging Face/OpenAI backends—typically coordinating generation, lightweight tagging/summarization, and embeddings across three models—and it also supports a server mode and integration with the VS Code Continue interface.

Li, Mengnan [Idaho National Laboratory (INL), Idah↗

Boundary-Aware Adversarial Learning Domain Adaption and Active Learning for Cross-Sensor Building Extraction

The use of convolutional neural networks (CNNs) for building extraction from remote sensing images has been widely studied and many public datasets have been made available for accelerating development of these CNN models. Yet adapting pretrained models at scale in real-world scenarios remains a challenging task. The main barrier is that certain new labels are still needed to compensate for domain shifting between the labeled data and new images that potentially cover new geographic locations or that are from a different sensor. In this article, we propose to add informatively labeled samples from a new image pool under the paradigm of active learning. To select the most useful samples based on model uncertainty, we first tackle the problem of uncalibrated uncertainty estimation due to distribution shifting by adapting feature extractors with boundary-based adversarial learning. Calibrated uncertainty is used as the query criterion in the active learning process, where the most uncertain samples are selected for annotation and included for model retraining. The proposed workflow was tested with three data pairs in which each workflow represents a scenario often encountered in real-world applications, including adapting pretrained models to new images collected with different sensors or to new geographic areas where appearances and types of buildings are very different. Compared to several baselines, including random sampling, temperature scaling (a well-known uncertainty calibration technique), different query strategies, and active domain adaptation methods, the proposed workflow shows that strategically querying a smaller set of samples for labeling achieves comparable or better building extraction performance. The proposed method reduces the number of labeled samples required to achieve sufficient model accuracy, thus significantly reducing hundreds of person-hours for labeled data creation. In addition, we include a few considerations when deploying this workflow in a GPU cluster that can be easily adapted to achieve operational building extraction model retraining.

97 MATHEMATICS AND COMPUTING↗

Robust Containment Queries over Collections of Trimmed NURBS Surfaces via Generalized Winding Numbers

Here, we propose a containment query that is robust to the watertightness of regions bound by trimmed NURBS surfaces, as this property is difficult to guarantee for in-the-wild CAD models. Containment is determined through the generalized winding number (GWN), a mathematical construction that is indifferent to the arrangement of surfaces in the shape. Applying contemporary techniques for the 3D GWN to trimmed NURBS surfaces requires some form of geometric discretization, introducing computational inefficiency to the algorithm and even risking containment misclassifications near the surface. In contrast, our proposed method leverages properties of the 3D solid angle to solve the relevant surface integral using a boundary formulation with rapidly converging adaptive quadrature. Batches of queries are further accelerated by memoizing (i.e., caching and reusing) quadrature node positions and tangents as they are evaluated. We demonstrate that our GWN method is robust to complex trimming geometry in a CAD model, and is accurate up to arbitrary precision at arbitrary distances from the surface. The derived containment query is therefore robust to model non-watertightness while respecting all curved features of the input shape.

97 MATHEMATICS AND COMPUTING↗

Scalable Volume Visualization for Big Scientific Data Modeled by Functional Approximation

Considering the challenges posed by the space and time complexities in handling extensive scientific volumetric data, various data representations have been developed for the analysis of large-scale scientific data. Multivariate functional approximation (MFA) is an innovative data model designed to tackle substantial challenges in scientific data analysis. It computes values and derivatives with high-order accuracy throughout the spatial domain, mitigating artifacts associated with zero- or first-order interpolation. However, the slow query time through MFA makes it less suitable for interactively visualizing a large MFA model. In this work, we develop the first scalable interactive volume visualization pipeline, MFA-DVV, for the MFA model encoded from large-scale datasets. Our method achieves low input latency through distributed architecture, and its performance can be further enhanced by utilizing a compressed MFA model while still maintaining a high-quality rendering result for scientific datasets. We conduct comprehensive experiments to show that MFA-DVV can decrease the input latency and achieve superior visualization results for big scientific data compared with existing approaches.

big scientific dataset↗

Evaluation of dual-weighted residual and machine learning error estimation for projection-based reduced-order models of steady partial differential equations

Projection-based reduced-order models (pROMs) show great promise as a means to accelerate many-query applications such as forward error propagation, solving inverse problems, and design optimization. In order to deploy pROMs in the context of high-consequence decision making, accurate error estimates are required to determine the region(s) of applicability in the parameter space. The following paper considers the dual-weighted residual (DWR) error estimate for pROMs and compares it to another promising pROM error estimate, machine learned error models (MLEM). Here, we show how DWR can be applied to ROMs and then evaluate DWR on two partial differential equations (PDEs): a two-dimensional linear convection–reaction–diffusion equation, and a three-dimensional static hyper-elastic beam. It is shown that DWR is able to estimate errors for pROMs extrapolating outside of their training set while MLEM is best suited for pROMs used to interpolate within the pROM training set.

42 ENGINEERING↗

BEAST DB: Grand-Canonical Database of Electrocatalyst Properties

We present BEAST DB, an open-source database comprised of ab initio electrochemical data computed using grand-canonical density functional theory in implicit solvent at consistent calculation parameters. The database contains over 20,000 surface calculations and covers a broad set of heterogeneous catalyst materials and electrochemical reactions. Calculations were performed at self-consistent fixed potential as well as constant charge to facilitate comparisons to the computational hydrogen electrode. This article presents common use cases of the database to rationalize trends in catalyst activity, screen catalyst material spaces, understand elementary mechanistic steps, analyze the electronic structure, and train machine learning models to predict higher fidelity properties. Users can interact graphically with the database by querying for individual calculations to gain a granular understanding of reaction steps or by querying for an entire reaction pathway on a given material using an interactive reaction pathway tool. BEAST DB will be periodically updated, with planned future updates to include advanced electronic structure data, surface speciation studies, and greater reaction coverage.

database↗

Investigating Kinetic Mechanisms of Soot Formation in Plasma Pyrolysis of Methane via Active Learning (Final Technical Report)

Plasma pyrolysis of methane is an effective route for zero-carbon hydrogen production. Yet, soot generated from pyrolysis of hydrocarbons is detrimental to the climate and human health. There is ample experimental and theoretical evidence that suggests polycyclic aromatic hydrocarbons (PAHs) are the molecular precursors to soot particles. The reaction pathways of PAH formation are intricately dependent on a multitude of process parameters, whose kinetic mechanisms are not well-understood in plasma pyrolysis. This project aims to leverage advances in the kinetic modeling of soot formation in combustion, as well as in surrogate modeling and active learning, to systematically investigate the effects of process parameter on the kinetics of PAH formation in plasma pyrolysis of methane. To this end, we propose to use the PAH formation kinetics model developed by the PPPL/PU group based on the well-established ABF and HACA mechanisms, coupled with low-temperature plasma models. We will develop an active learning (AL) framework based on Bayesian optimization to systematically and data-efficiently explore the complex and multivariable parameter space of plasma pyrolysis in order to quantify the effects of plasma and feed parameters on the ABF and HACA kinetic pathways. AL is the branch of machine learning concerned with systematically querying samples from a system (experimental or computational) to train a data-driven model that maps design parameters to a performance criterion. We will use the data generated via AL to perform global sensitivity analysis, combined with uncertainty quantification, to elucidate the impact of different reaction pathways on minimizing formation of soot precursors. This study will result in an improved understanding of kinetics of PAH formation in plasma pyrolysis and can pave the way for more advanced mechanistic studies (e.g., soot nucleation mechanisms). Additionally, the findings will be useful for establishing practical strategies for increasing the pyrolysis efficiency and producing high-grade carbon for synthesis of nanomaterials.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

CAMFeND: Credibility-Aware Multimodal Fake News Detection with Rotational Attention

In the evolving digital landscape, fake news is a significant challenge, influencing public perception and decision-making. Traditional detection approaches focus on single-modal data or simple multimodal fusion, often overlooking deeper interactions and news credibility. We propose a novel model addressing these limitations by introducing rotational attention and news domain information as a feature. Unlike static attention mechanisms, our rotational attention dynamically shifts query, key, and value roles across text and image inputs, enabling richer cross-modal interaction. Incorporating news domain information further enhances the model’s reliability by associating news posts with top domains extracted from Google search results, reducing false detections. This approach assesses both the content and the broader web context in which the news is discussed. Our model outperforms existing state-of-the-art methods by providing deeper, layered multimodal integration and domain information analysis, resulting in a more robust and adaptive fake news detection system.

Gupta, Nidhi↗

Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning

Abstract Motivation Creating knowledge bases and ontologies is a time consuming task that relies on manual curation. AI/NLP approaches can assist expert curators in populating these knowledge bases, but current approaches rely on extensive training data, and are not able to populate arbitrarily complex nested knowledge schemas. Results Here we present Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES), a Knowledge Extraction approach that relies on the ability of Large Language Models (LLMs) to perform zero-shot learning and general-purpose query answering from flexible prompts and return information conforming to a specified schema. Given a detailed, user-defined knowledge schema and an input text, SPIRES recursively performs prompt interrogation against an LLM to obtain a set of responses matching the provided schema. SPIRES uses existing ontologies and vocabularies to provide identifiers for matched elements. We present examples of applying SPIRES in different domains, including extraction of food recipes, multi-species cellular signaling pathways, disease treatments, multi-step drug mechanisms, and chemical to disease relationships. Current SPIRES accuracy is comparable to the mid-range of existing Relation Extraction methods, but greatly surpasses an LLM’s native capability of grounding entities with unique identifiers. SPIRES has the advantage of easy customization, flexibility, and, crucially, the ability to perform new tasks in the absence of any new training data. This method supports a general strategy of leveraging the language interpreting capabilities of LLMs to assemble knowledge bases, assisting manual knowledge curation and acquisition while supporting validation with publicly-available databases and ontologies external to the LLM. Availability and implementation SPIRES is available as part of the open source OntoGPT package: https://github.com/monarch-initiative/ontogpt.

59 BASIC BIOLOGICAL SCIENCES↗

Direct Numerical Simulation Database of High-Speed Flow over Parameterized Curved Walls

This study presents a direct numerical simulation (DNS) database of high-speed turbulent boundary layers (TBLs) subject to pressure gradients due to parametrically varied backward-facing and forward-facing wall curvatures, with an inflow Mach number of 4.9 and a friction Reynolds number of [Formula: see text] immediately before the onset of wall curvature. The Mach and Reynolds numbers are significantly higher than those reported in the literature for the DNS of pressure-gradient TBLs. The flow conditions and baseline wall geometries are representative of experiments in the high-speed blowdown wind tunnel at the National Aerothermochemistry Laboratory at Texas A&M University. The wall steepness of the baseline geometry for both the backward-facing and forward-facing walls was systematically varied to cause attached, incipiently separated, and fully separated flows. Precomputed flow statistics, including turbulent kinetic energy budgets, are available on the website of the Turbulence Modeling Resource of the NASA Langley Research Center, allowing other investigators to query any property of interest.

Engineering↗

LAF-Net: A Deep Residual and Cross-Attention Framework for Day-Ahead Load Forecasting: Preprint

Accurate day-ahead load forecasting is essential for reliable power system operations and market efficiency. System operators such as the Midcontinent Independent System Operator (MISO) rely on forecasts from multiple vendors, yet combining them effectively remains a persistent challenge due to vendor-specific biases. This paper presents a novel LSTM-Attention Fusion Network with Error Representation (LAF-Net) that enhances day-ahead hourly load forecasting through deep residual learning and multi-modal cross-attention. The proposed model builds a historical error memory from past vendor performance and dynamically queries it with future hour context to generate adaptive, hour-specific trust weights for each vendor. A bounded residual correction further refines forecasts by mitigating systematic and temporally localized errors. Tested on real MISO LBA data with multi-vendor forecasts, LAF-Net consistently outperforms the best vendor baseline across all 38 LBAs, achieving more than a 40% reduction in system-level mean absolute error (MAE) during peak load hours relative to the best vendor baseline.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Do-calculus enables estimation of causal effects in partially observed biomolecular pathways

Abstract Motivation Estimating causal queries, such as changes in protein abundance in response to a perturbation, is a fundamental task in the analysis of biomolecular pathways. The estimation requires experimental measurements on the pathway components. However, in practice many pathway components are left unobserved (latent) because they are either unknown, or difficult to measure. Latent variable models (LVMs) are well-suited for such estimation. Unfortunately, LVM-based estimation of causal queries can be inaccurate when parameters of the latent variables are not uniquely identified, or when the number of latent variables is misspecified. This has limited the use of LVMs for causal inference in biomolecular pathways. Results In this article, we propose a general and practical approach for LVM-based estimation of causal queries. We prove that, despite the challenges above, LVM-based estimators of causal queries are accurate if the queries are identifiable according to Pearl’s do-calculus and describe an algorithm for its estimation. We illustrate the breadth and the practical utility of this approach for estimating causal queries in four synthetic and two experimental case studies, where structures of biomolecular pathways challenge the existing methods for causal query estimation. Availability and implementation The code and the data documenting all the case studies are available at https://github.com/srtaheri/LVMwithDoCalculus. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Consist v0.1.0

A Python library for provenance tracking, intelligent caching, and data virtualization in scientific simulation workflows. It automatically records code, configuration, and input data to skip redundant computations and enables querying results across many runs without manual bookkeeping. Designed to support multi-model simulation workflows like the BEAM CORE toolset at LBL, but designed to be extensible to a wide range of research workflows. Combines lineage tracking features as provided by OpenLineage with deterministic hashing like SnakeMake, and adds powerful analysis tools on model outputs.

Needell, Zachary [Lawrence Berkeley National Labor↗