Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “information retrieval”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Self-supervised physics-informed generative networks for phase retrieval from a single X-ray hologram

X-ray phase contrast imaging significantly improves the visualization of structures with weak or uniform absorption, broadening its applications across a wide range of scientific disciplines. Propagation-based phase contrast is particularly suitable for time- or dose-critical in vivo/in situ/operando (tomography) experiments because it requires only a single intensity measurement. However, the phase information of the wave field is lost during the measurement and must be recovered. Conventional algebraic and iterative methods often rely on specific approximations or boundary conditions that may not be met by many samples or experimental setups. In addition, they require manual tuning of reconstruction parameters by experts, making them less adaptable for complex or variable conditions. Here we present a self-learning approach for solving the inverse problem of phase retrieval in the near-field regime of Fresnel theory using a single intensity measurement (hologram). A physics-informed generative adversarial network is employed to reconstruct both the phase and absorbance of the unpropagated wave field in the sample plane from a single hologram. Unlike most state-of-the-art deep learning approaches for phase retrieval, our approach does not require paired, unpaired, or simulated training data. This significantly broadens the applicability of our approach, as acquiring or generating suitable training data remains a major challenge due to the wide variability in sample types and experimental configurations. The algorithm demonstrates robust and consistent performance across diverse imaging conditions and sample types, delivering quantitative, high-quality reconstructions for both simulated data and experimental datasets acquired at beamline P05 at PETRA III (DESY, Hamburg), operated by Helmholtz-Zentrum Hereon. Furthermore, it enables the simultaneous retrieval of both phase and absorption information.

36 MATERIALS SCIENCE↗

Roadmap on data-centric materials science

Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

36 MATERIALS SCIENCE↗

Relaxing Direct Ptychography Sampling Requirements via Parallax Imaging Insights

Direct ptychography enables the retrieval of information encoded in the phase of an electron wave passing through a thin sample by deconvolving the interference effects of a converged probe with known aberrations. Under the weak phase object approximation, this permits the optimal transfer of information using noniterative techniques. However, the achievable resolution of the technique is traditionally limited by the probe step size—setting stringent Nyquist sampling requirements. At the same time, parallax imaging has emerged as a dose-efficient phase technique which relaxes sampling requirements and enables scan-upsampling. Here, we formulate parallax imaging as a quadratic approximation to part of the direct ptychography kernel and use this insight to enable upsampling in direct ptychography. We validate our analytical results numerically using simulated and experimental reconstructions.

direct ptychography, parallax imaging, scan Nyquis↗

Replacing non-biomedical concepts improves embedding of biomedical concepts

Embeddings are semantically meaningful representations of words in a vector space, commonly used to enhance downstream machine learning applications. Traditional biomedical embedding techniques often replace all synonymous words representing biological or medical concepts with a unique token, ensuring consistent representation and improving embedding quality. However, the potential impact of replacing non-biomedical concept synonyms has received less attention. Embedding approaches often employ concept replacement to replace concepts that span multiple words, such as non-small-cell lung carcinoma, with a single concept identifier (e.g., D002289). Also, all synonyms of each concept are merged into the same identifier. Here, we additionally leveraged WordNet to identify and replace sets of non-biomedical synonyms with their most common representatives. This combined approach aimed to reduce embedding noise from non-biomedical terms while preserving the integrity of biomedical concept representations. We applied this method to 1,055 biomedical concept sets representing molecular signatures or medical categories and assessed the mean pairwise distance of embeddings with and without non-biomedical synonym replacement. A smaller mean pairwise distance was interpreted as greater intra-cluster coherence and higher embedding quality. Embeddings were generated using the Word2Vec algorithm applied to a corpus of 10 million PubMed abstracts. Our results demonstrate that the addition of non-biomedical synonym replacement reduced the mean intra-cluster distance by an average of 8%, suggesting that this complementary approach enhances embedding quality. Future work will assess its applicability to other embedding techniques and downstream tasks. Python code implementing this method is provided under an open-source license.

algorithms↗

2002 Treasure Valley Transportation Survey

The 2002 Treasure Valley Transportation Survey was conducted in Ada and Canyon counties in southwest Idaho, under contract with the Community Planning Association of Southwest Idaho. The full study was conducted during September 2002 and October 2002 and entailed the collection of activity and travel information for all household members, regardless of age, during an assigned 24-hour period (Tuesday, Wednesday, or Thursday). In addition to providing basic demographic information about each household and its members, the survey documented specific travel characteristics and trips made, including the number of occupants, trip purpose, time of day, and questions specific to mode use. Travel days for the survey were spread across the pilot study (August 8, 2002) and the full study (September 3, 2002-October 31, 2002). In total, 3,488 households were recruited to participate in the study. Of these, 2,582 completed travel diaries (fully completed and passed edit check procedures), and the information was retrieved from all household members.

1Hz data↗

2001-2003 Ohio Statewide Household Travel Survey

The purpose of this study, conducted under the auspices of the Ohio Department of Transportation between August 2001 and May 2003, was to update the statewide database of household socioeconomic and travel information. Data insights were used to refine travel estimates, models, and forecasts throughout Ohio and specifically for the nine smallest metropolitan planning organizations (MPOs). The study area consists of all Ohio counties except for those that are within the MPO boundaries of Cincinnati, Cleveland, and Columbus. The study is an essential element in determining statewide and regional travel patterns. A total of 16,112 households provided recruitment and travel/activity data, and the information was retrieved from all household members regardless of age. A total of 122,463 trips were profiled during the survey during the 24-hour assigned weekday.

1Hz data↗

1995 San Diego Region Travel Behavior Survey

The 1995 San Diego Region Travel Behavior Survey was conducted January-June of 1995 under the auspices of the San Diego Association of Governments. The survey was an essential element in the regional study of transportation activity and travel patterns. The survey was designed to examine the relationship between characteristics of households and travel behavior, as well as to provide information state and local decision makers require when considering future regional transportation needs and investments. In total, 2,375 households were recruited to participate in the study. Of these, 2,062 households completed travel diaries, and the information was retrieved from all household members older than age five. 2,049 were successfully geocoded to California state plane coordinates for home addresses. A total of 17,060 trips were profiled in the course of this survey during the assigned 24-hour weekday.

1Hz data↗

1999 Puget Sound Household Travel Survey

The 1999 Puget Sound Household Travel Survey was conducted between July and November. NuStats Research and Consulting conducted the survey on behalf of the Puget Sound Regional Council. The purpose of the study was to provide data to continue developing and refining the Regional Travel Demand Forecasting Model, as well as to provide a better understanding of travel behavior in the Puget Sound region. The study area consists of King, Kitsap, Pierce, and Snohomish counties. The resultant dataset will be used to fulfill the model's functions of estimating trip generation and distribution, mode choice, and assignments. The study had household members 16 years old or older keep track of travel for a 48-hour period. A total of 9,028 households were recruited to participate in the study. Of these, 6,000 households (66.5%) completed travel diaries, and the information was retrieved from all household members regardless of age. An “attitude” survey about transportation and land use issues was also mailed to household members 16 years old or older.

1Hz data↗

1998/99 Thurston County Household Travel Study

The survey was conducted under the auspices of the Thurston Regional Planning Council, and it was funded through a state grant awarded to Intercity Transit of Olympia, Washington. Data collection was from September 1998 through March 1999. The purpose of the study was to provide data for the continuing development and refinement of the Regional Travel Demand Forecasting Model, as well as to provide a better understanding of travel behavior in the southern Puget Sound region of Washington. The resultant data set will be used to fulfill the model's functions of estimating trip generation and distribution, mode choice, and assignments. Participating households were assigned specific “travel days” to record their travel over a 48-hour period. A total of 2,465 households were recruited to participate in the study. Of these, 1,537 households completed travel diaries, and the information was retrieved from 3,653 household members regardless of age. Households member made 25,278 total trips during their 48-hour diary period.

1Hz data↗

Using a Large Language Model for Accurate Technical Language Generation in the Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants

Machine learning (ML) methods for predictive maintenance (PdM) are emerging as effective proactive strategies for diagnosing equipment degradation and enabling effective decision-making. However, explainability and trustworthiness of artificial intelligence are two salient challenges that need to be addressed for wider deployment of these technologies in nuclear power plants (NPPs). Large language models (LLMs) offer a unique approach to tackle these challenges by explaining PdM, work orders, diagnosis results, and ML algorithms to users, who may not be familiar with ML and PdM in general. Moreover, by dynamically retrieving relevant information from technical documents and evaluating factuality of LLM generation, the accuracy and relevance of LLM generations can be improved. This work demonstrates using LLMs to explain the causes and consequences of circulating water system failures based on multiyear NPP work orders. This work tests the capability of multimodal LLM approaches in explaining the differences in the circulating water system from both the Salem and Hope Creek NPPs using both text and image resources. This work also demonstrates the use of multimodal LLMs in describing the diagnosis tab of a predictive maintenance software named VIsualization for PrEdictive maintenance Recommendation (VIPER) to users.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Mnemosyne

SAND2024-08570O Mnemosyne is an interactive tool for finding, retrieving, and exploring information about U.S. nuclear tests documented in the National Nuclear Security Administration’s NV-209 report. It also acts as an information architecture and codebase for integrating additional information and computational tools related to these tests at the unclassified and classified levels. Users can search by any number of nuclear test attributes—name, yield range, altitude ranges, purpose of a test—and find all matching tests. Users can also retrieve specific test information published in NV-209. The tool displays geospatial and topological data about test location, and it provides an information architecture for storing additional contextual material, such as photographs. Mnemosyne provides a capability of interfacing with HYCHEM, Sandia's nuclear detonation optical waveform tool. The software is designed for use by government, academia, military, and research institutions. The software will likely be advanced to integrate seismic data and the nuclear detonation optical signal simulation code radCTH. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Fisher, Dustin↗

Six-Letter DNA Nanotechnology: Incorporation of Z-P Base Pairs into Self-Assembling 3D Crystals

Artificially expanded genetic information systems (AEGIS) were developed to expand the diversity and functionality of biological systems. Recent experiments have shown that these expanded DNA molecular systems are robust platforms for information storage and retrieval as well as useful for basic biotechnologies. In tandem, nucleic acid nanotechnology has seen the use of information-based “semantomorphic” encoding to drive the self-assembly of a vast array of supramolecular devices. To establish the effectiveness of AEGIS toward nanotechnological applications, we investigated the ability of a six-letter alphabet composed of A:T, G:C and synthetic Z:P (Z, 6-amino-3-(1'-β- D-2'-deoxy ribofuranosyl)-5-nitro-(1H)-pyridin-2-one; P, 2-amino-8-(1'- β-D-2'-deoxyribofuranosyl)-imidazo-[1,2a]-1,3,5-triazin-(8H)-4-one) base pairs to engage in 3D self-assembly. We found that crystals could be programmably assembled from AEGIS oligomers. We conclude that unnatural base pairs can be used for the topological self-assembly of crystals. We anticipate the expansion of AEGISbased nucleic acid nanotechnologies to enable the development of novel nanomaterials, high-fidelity signal cascades, and dynamic nanoscale devices.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluating the Effectiveness of Retrieval-Augmented Large Language Models in Scientific Document Reasoning

Despite the dramatic progress in Large Language Model (LLM) development, LLMs often provide seemingly plausible but not factual information, often referred as hallucinations. Retrieval-augmented LLMs provide a non-parametric approach to solve these issues by retrieving relevant information from external data sources and augment the training process. These models helps to trace evidence from an externally provided knowledge base allowing the model predictions to be better interpreted and verified. In this work, we critically evaluate these models in their ability to perform in scientific document reasoning tasks. To this end, we tuned multiple such model variants with science-focused instructions and evaluated them on a scientific document reasoning benchmark for the usefulness of the retrieved document passages. Our findings suggest that models justify predictions in science tasks with fabricated evidence and leveraging scientific corpus as pretraining data does not alleviate the risk of evidence fabrication.

• Artificial intelligence (AI) / machine learning ↗

Emergent Dimer-Model Topological Order and Quasiparticle Excitations in Liquid Crystals: Combinatorial Vortex Lattices

Liquid crystals have proven to provide a versatile experimental and theoretical platform for studying topological objects such as vortices, skyrmions, and hopfions. In parallel, in hard condensed matter physics, the concept of topological phases and topological order has been introduced in the context of spin liquids to investigate emergent phenomena like quantum Hall effects and high-temperature superconductivity. Here, we bridge these two seemingly disparate perspectives on topology in physics. Combining experiments and simulations, we show how topological defects in liquid crystals can be used as versatile building blocks to create complex, highly degenerate topological phases, which we refer to as “combinatorial vortex lattices” (CVLs). CVLs exhibit extensive residual entropy and support locally stable quasiparticle excitations in the form of charge-conserving topological monopoles, which can act as mobile information carriers and be linked via Dirac strings. CVLs can be rewritten and reconfigured on demand, endowed with various symmetries, and modified through laser-induced topological surgery—an essential capability for information storage and retrieval. We demonstrate experimentally the realization, stability, and precise optical manipulation of CVLs, thus opening new avenues for understanding and technologically exploiting higher-hierarchy topology in liquid crystals and other ordered media.

36 MATERIALS SCIENCE↗

CHESS 2025: Orthorectified airborne RGB imagery from NEON AOP surveys

This dataset provides Level 1 (L1) and Level 3 (L3) orthorectified Red-Green-Blue (RGB) imagery collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). This high-resolution imagery is a photographic record of red, green, and blue visible light from sunlight reflected off of the Earth’s surface. The data comprise full-color images of the ground surface and are primarily intended to provide context to imaging spectroscopy and light detection and ranging (LiDAR) data. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. RGB images were acquired using the PhaseOne IXM-RS150F high-resolution digital camera onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). The package data include both an L1 product comprising one camera frame per file and an L3 mosaic aligned to the Universal Transverse Mercator (UTM) Zone 13N grid and the World Geodetic System (WGS) 84 projection. Both products are provided in geotif (.tif) format at 0.1 m ground resolution. The bulk of the imagery was collected during the main CHESS field campaign from June 13 to July 15, 2025. Additional images of a portion of the Upper Taylor (UPTA) domain were collected on September 18, 2025, to fill gaps in imagery identified after the main campaign was complete. RGB camera imagery is not radiometrically calibrated, and therefore pixel values should not be exploited for scientific analysis. Pixel values have undergone a manual adjustment to enhance feature identification. The imagery is rigorously geolocated which does allow for reliable geometric information to be retrieved. To generate the orthorectified imagery, the NEON AOP camera captured visible spectrum in red, green, and blue bands. The raw images were then processed using NEON’s camera orthorectification workflow. A boresight calibration flight was made to build a complete camera, distortion, and alignment model. Color balance/white balance and exposure correction were applied to the raw RGB images. The corrected images were orthorectified by ray-tracing image pixels to a lidar-derived digital surface model (DSM) mesh using the refined camera model, outputting orthorectified raster pixels on a regular grid. Flightline-level data were mosaicked by selecting per-pixel contributions from overlapping orthorectified images using line-of-sight (LOS) zenith angle minimization to reduce edge distortions. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Development and Implementation of a New AI-Based Tool to Support Fast Reactor Software Model Generation and Validation

This report summarizes FY26 work to develop Maggie, an artificial intelligence-based assistant designed to support software model generation and validation activities for fast reactor analysis codes. The project established a modular, code-agnostic software architecture that separates reusable agent capabilities from code-specific knowledge and tools, with initial implementation focused on the FRP-supported fast reactor safety analysis code SAS4A/SASSYS1 (SAS). A curated SAS-specific knowledge base was assembled from the code manual, training materials, historical analysis reports, and representative input files, and was integrated through retrieval-augmented generation to ground Maggie’s responses in authoritative sources. Maggie was deployed on the internal Argonne network, where it demonstrated practical user-facing capability as a chatbot for answering natural language questions about SAS and retrieving relevant technical information. Demonstration cases also showed that Maggie can generate useful snippets of SAS input for selected modeling tasks, while highlighting current limitations in reliability and consistency for more complex input generation tasks. Overall, the FY26 effort established the technical foundation for an AI-assisted capability intended to improve the efficiency, consistency, and accessibility of fast reactor software model development at Argonne and, with further improvements, to support eventual use by the broader fast reactor community, including industry users of FRP-supported analysis tools.

Thomas, Rachel [Argonne National Laboratory (ANL),↗

Quality Control of Silicon Sensor Modules for Particle Detectors

The High-Luminosity Large Hadron Collider (HL-LHC) will produce a higher rate of particle collisions than the current Large Hadron Collider (LHC), requiring important upgrades to the Compact Muon Solenoid (CMS) to handle an increased amount of data. An important upgrade is the Phase-2 Outer Tracker Upgrade, which consists of 13,000 silicon sensor modules made of two parallel silicon sensors and readout electronics. These modules undergo careful quality control checks both during and after module assembly to ensure precise and reliable detector performance. This project focuses on precision testing for quality control of silicon sensor modules at Fermilab. Hands-on work includes visual inspection, current-voltage testing, module testing, and ultraviolet (UV) light exposure of modules showing abnormal current-voltage behavior. The ultraviolet exposure process improves the abnormal sensor readout data by placing the selected sensor side of the module directly under the UV light inside a controlled box. In addition to laboratory testing and ultraviolet experiments, I developed a Python-based data tool that connects to a module database and allows selected testing conditions and module information to be retrieved and displayed efficiently. These different testing procedures, experimental processes, and computational tools support the broader goal of identifying module issues and improving modules that will be used in the CMS Outer Tracker Phase-2 Upgrade.

Siddiqui, Hooriya [DuPage Coll.] (ORCID:0009000151↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on ~30 m range gates, stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, below range, ran out of signal, cloud-topped). Cloud Base Height (Haar-gradient detection): 15 min estimates of cloud-base height (m) with a cloud-detection quality flag (0–3: none, low, moderate, high). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (2.0.0), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution, with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗