Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “text analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Leveraging BERT and Network-Based Attention Analysis for Identifying Treatment Milestones in EHRs

This study introduces a sophisticated data-driven framework for analyzing Electronic Health Records (EHRs) using transformer-based models to identify and disentangle overlapping treatment contexts. The framework leverages a preprocessing pipeline that transforms structured procedural codes into semantically enriched descriptive text, enabling the use of attention mechanisms to cluster medical events into treatment milestones—cohesive and distinct components of care processes. The methodology is rigorously validated using synthetic datasets derived from the MIMIC-III database, designed to simulate the heterogeneity and overlapping procedural contexts characteristic of real-world EHR scenarios. Quantitative evaluation highlights the framework’s robustness in disentangling concurrent care pathways, with attention metrics and unsupervised clustering approaches demonstrating the ability to preserve intra-context relationships while distinguishing inter-context dependencies. By addressing challenges inherent in data heterogeneity, this approach provides a foundation for uncovering complex treatment patterns, advancing clinical decision-making, and optimizing resource allocation in diverse healthcare environments.

Kim, Minsu [ORNL] (ORCID:0000000224185535)↗

Theorems in Service of Sound Composition, Rapid Modeling and Scalable Analysis

This project extends the state of the art in formal verification modeling with modules and automatically checkable data-sharing patterns such that component modules can retain their assurance case when composed within a larger system. For users, smaller models make reasoning easier and help to ensure they accurately reflect text specifications. For automated methods, smaller models give exponential benefits for verification algorithm execution time.

97 MATHEMATICS AND COMPUTING↗

45Q Addendum to the NETL CO2U LCA Guidance Document (V.2.0)

This document provides additional guidance and changes to the Carbon Dioxide Utilization Life Cycle Analysis Guidance for the U.S. DOE Office of Fossil Energy and Carbon Management, Version 2.0 to make it more applicable to taxpayers preparing life cycle analyses for the 45Q tax credit. This update adds clarity to existing text and includes information on new tools and data available.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

FolpsD: combining EFT and phenomenological approaches for joint power spectrum and bispectrum analyses

We present a theoretical model for the power spectrum and bispectrum of galaxy clustering that exploits the complementarity between small-scale power spectrum information and large-scale bispectrum measurements. We extend the FOLPS code by combining its one-loop EFT galaxy power spectrum with a tree-level galaxy bispectrum projected onto the tripolar spherical harmonics (Sugiyama) basis. To access additional small-scale information, we also consider a line-of-sight damping factor in both statistics, mirroring approaches commonly used in studies of redshift-space distortions. We test the model using DESI DR2 galaxy mocks. Even without damping, the joint analysis of the EFT power spectrum and bispectrum significantly improves constraints and reduces parameter degeneracies relative to power spectrum analyses alone. For LRG-like samples, including the damping further extends the range beyond $k\sim 0.3 \,h \text{Mpc}^{-1}$ in the power spectrum and $k \sim 0.24 \,h \text{Mpc}^{-1}$ in the bispectrum without introducing statistically significant parameter biases. This leads to up to $\sim 30\%$ tighter constraints on $A_s$ and $ω_{cdm}$. For low signal-to-noise tracers such as QSOs, however, the damping parameters are weakly constrained and can absorb noise fluctuations, leading to shifts in inferred parameters. Similar limitations may arise in models where cosmological information is encoded in power-spectrum shape features degenerate with the damping, such as scenarios with massive neutrinos. In contrast, for $w_0w_a$CDM we obtain $15\%$ and $21\%$ tighter constraints on $w_0$ and $w_a$, respectively, yielding a deviation from constant dark energy at slightly more than the $1σ$ level using full-shape information alone. The code is publicly available at https://github.com/cosmodesi/FolpsD

Bansal, P. [Michigan U., MCTP; Michigan U.] (ORCID↗

LLM integration into EPICS

The utilization of large language models (LLMs) such as ChatGPT has seen a remarkable increase in various fields over the past few years. These models have demonstrated their versatility and capability in understanding and generating human-like text, making them invaluable tools in numerous applications. In this project, we explore the integration of a LLM into the Experimental Physics and Industrial Control System (EPICS). The primary focus of this integration is to employ the LLM for advanced image processing and spatial analysis on images obtained from the beamlines. By leveraging the capabilities of the LLM, we aim to enhance the accuracy and efficiency of image interpretation, enabling more precise data analysis and decision-making within the EPICS framework. This integration not only showcases the potential of LLMs in scientific and industrial applications but also sets the stage for future advancements in automated control systems.

Adams, Ethan↗

Search for excited tau leptons in the ττγ final state in proton-proton collisions at $\sqrt{\text{s}}$ = 13 TeV

Results are presented for a test of the compositeness of the heaviest charged lepton, τ, using data collected by the CMS experiment in proton-proton collisions at a center-of-mass energy of 13 TeV at the CERN LHC. The data were collected in 2016–2018 and correspond to an integrated luminosity of 138 fb −1 . This analysis searches for tau lepton pair production in which one of the tau leptons is produced in an excited state and decays to a ground state tau lepton and a photon. The event selection consists of two isolated tau lepton decay candidates and a high-energy photon. The mass of the excited tau lepton is reconstructed using the missing transverse momentum in the event, assuming the momentum of the neutrinos from each tau lepton decay are aligned with the visible decay products. No excess of events above the standard model background prediction is observed. This null result is used to set lower bounds on the excited tau lepton mass. For a compositeness scale Λ equal to the excited tau lepton mass, excited tau leptons with masses below 4700 GeV are excluded at 95% confidence level; for Λ = 10 TeV this exclusion is set at 2800 GeV. This is the first experimental result covering this production and decay process in the excited tau mass range above 175 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Search for vector-like leptons with long-lived particle decays in the CMS muon system in proton-proton collisions at $$\sqrt{\text{s}}$$ = 13 TeV

Abstract A first search is presented for vector-like leptons (VLLs) exclusively decaying into a light long-lived pseudoscalar boson and a standard model τ lepton. The pseudoscalar boson is assumed to have a mass below the τ + τ − threshold, so that it decays exclusively into two photons. It is identified using the CMS muon system. The analysis is carried out using a data set of proton-proton collisions at a center-of-mass energy of 13 TeV collected by the CMS experiment in 2016–2018, corresponding to an integrated luminosity of 138 fb −1. Selected events contain at least one pseudoscalar boson decaying electromagnetically in the muon system and at least one hadronically decaying τ lepton. No significant excess of data events is observed compared to the background expectation. Upper limits are set at 95% confidence level on the vector-like lepton production cross section as a function of the VLL mass and the pseudoscalar boson mean proper decay length. The observed and expected exclusion ranges of the VLL mass extend up to 700 and 670 GeV, respectively, depending on the pseudoscalar boson lifetime.

Chekhovsky, V. [Yerevan Physics Institute]↗

First results from the search of $ν_μ$ disappearance with ICARUS

his work presents a search for muon neutrino disappearance signal using the ICARUS detector. ICARUS is currently working as the far detector of the Short-Baseline Neutrino (SBN) program. Its main goal is to investigate the possible oscillation of a sterile neutrino state ($\Delta m^2 \sim \text{eV}^2$), that would drive short baseline oscillations at the scale of L/E $\sim$ km/GeV. Neutrino interactions from the Booster Neutrino Beam (BNB) are investigated, looking for simple final state topology with a single muon and at least one proton (1$\mu$Np). This analysis tests the 3+1 model using the two flavor neutrino approximation, where careful consideration has been given to the systematic uncertainties arising from flux, neutrino interaction, and detector models. No statistically significant $\nu_\mu$ disappearance is observed, hence 90\% C.L. exclusion contours are reported in the $\Delta m^2_{41} - \sin^22\theta_{\mu\mu}$ parameter space.

Artero Pons, Maria [INFN, Padua; U. Padua (main)]↗

SUNBIRD : a simulation-based model for full-shape density-split clustering

Combining galaxy clustering information from regions of different environmental densities can help break cosmological parameter degeneracies and access non-Gaussian information from the density field that is not readily captured by the standard two-point correlation function (2PCF) analyses. However, modelling these density-dependent statistics down to the non-linear regime has so far remained challenging. We present a simulation-based model that is able to capture the cosmological dependence of the full shape of the density-split clustering (DSC) statistics down to intra-halo scales. Our models are based on neural-network emulators that are trained on high-fidelity mock galaxy catalogues within an extended-ΛCDM framework, incorporating the effects of redshift-space, Alcock–Paczynski distortions, and models of the halo–galaxy connection. Our models reach sub-percent level accuracy down to $1 \, h^{-1}\text{Mpc}$ and are robust against different choices of galaxy–halo connection modelling. When combined with the galaxy 2PCF, DSC can tighten the constraints on ω cdm , σ 8 , and n s by factors of 2.9, 1.9, and 2.1, respectively, compared to a 2PCF-only analysis. DSC additionally puts strong constraints on environment-based assembly bias parameters.

79 ASTRONOMY AND ASTROPHYSICS↗

Mountain Basin Controls on the Snow-to-Streamflow Signal: An AIC-Weighted Multiple Linear Regression Framework

A regression-based analysis quantifies how basin characteristics modulate the snow-to-streamflow signal. First, we use the ERA5-Land reanalysis gridded product (European Centre for Medium Range Weather Forecasts reanalysis 5 -Land component) for 4,655 hydrologic unit code - 10 (HUC10) mountain basins across the western United States (US) for water years 1987–2024. Linear regressions are performed for peak snow water equivalent (SWE) and annual streamflow for each mountain basin. Models use ordinary least squares in Python’s statsmodels package. After which, an Akaike Information Criterion (AIC)–weighted ensemble multiple linear regression (MLR) framework with 47 watershed traits is used to predict the linear regression coefficient of determination (r-squared) defining the ability of peak SWE to predict annual streamflow across all mountain basin. Predictor sets are constrained to avoid multicollinearity by excluding models with variance inflation factors (VIF) greater than 5. Mountain basin traits included in the MLR include seasonal climate, topography, vegetation type and structure, and bedrock geology. Accepted models are considered if their AIC is within 2.0 of the model with the minimum AIC, or best model. To compare predictor influence across acceptable models, we computed standardized regression coefficients. To evaluate structural redundancy among models, we constructed binary inclusion vectors for each acceptable model, denoting whether a predictor was present (1) or absent (0). Core predictor variables are defined as occurring in at least 67% of the acceptable models. For this regional analysis, only one model was found acceptable, with higher snow-to-streamflow translation (higher r-squared) occurring in colder mountain basins with higher relative winter precipitation, more snow accumulation and a lower fraction of annual precipitation that falls in the spring and summer. The second component of the data package uses previously published, high-resolution output from an integrated hydrological model of the East River watershed using the U.S. Geological Survey Groundwater and Surface water Flow model (GSFLOW, doi:10.15485/1998576). East River MLR expands upon the approach described above to explore the response of five streamflow metrics—annual streamflow, runoff efficiency, 7-day minimum flow, low-flow duration, and non-perennial stream fraction to snow system indicators including peak SWE, snow-covered area, snow disappearance date, and the fraction of basin area characterized by low-to-no snow, as well as seasonal precipitation and temperature, and annual hydrologic variables representing soil moisture, evapotranspiration (ET), the partitioning of incoming precipitation to evapotranspiration (ET/P), groundwater storage, and groundwater inflow to streams. MLR was done on all water years (P0: 1987-2024) and for each period as determined in the split analysis using pooled regression techniques (P1: 1987-2011 and P2: 2012-2024) to evaluate shifting predictor variable emphasis on streamflow generation. Results indicate that since 2012, peak SWE has lost statistical strength in its prediction of annual streamflow and runoff efficiency, and the indirect influence of spring temperature has emerged as critically important. Low-flow metrics remain largely influenced by soil moisture, vegetation water use and groundwater inflows with summer precipitation becoming a direct influence on minimum summer flow. Together, these data and Python-based analysis tools provide a framework for identifying the key watershed characteristics that control how streamflow responds to snow from year to year. The package also helps quantify uncertainty in statistical models and assess how snow–streamflow relationships vary across regions and over time. This dataset contains comma-separated values files (.csv), text files (.txt), python code files (.py), figure files (.png), and shapefiles (.cpg, .dbf, .prj, .sbn, .sbx, .shp, .xml). Further details on file contents and MLR execution can be found in the readme file and the FLMD files. Work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

$\overline{\Sigma }^{\pm }$ production in $\text {pp}$ and $\text {p}{-}\text{Pb}$ collisions at $\sqrt{s_{\textrm{NN}}} = 5.02$ TeV with ALICE

The transverse momentum spectra and integrated yields of anti-$Σ$ hyperons ($\overline{\Sigma}^{\pm}$) have been measured in and collisions at $\sqrt{s_{\textrm{NN}}} = 5.02$ TeV with the ALICE experiment. Measurements are performed via the newly accessed decay channel $\overline{\Sigma}^{\pm}$ → $\bar{\textrm{n}}π^±$. A new method of antineutron reconstruction with the PHOS electromagnetic spectrometer is developed and applied to this analysis. The p T spectra of $\overline{\Sigma}^{\pm}$ are measured in the range 0.5 < p T < 3 GeV/c and compared to predictions of the PYTHIA 8, DPMJET, PHOJET, EPOS LHC and EPOS4 models. The EPOS LHC and EPOS4 models provide the best descriptions of the measured spectra both in pp and p-Pb collisions, while models which do not account for multiparton interactions provide a considerably worse description at high p T . The total yields of $\overline{\Sigma }^{\pm }$ in both pp and p-Pb collisions are compared to predictions of the Thermal-FIST model and dynamical models PYTHIA 8, DPMJET, PHOJET, EPOS LHC and EPOS4. All models reproduce the total yields in both colliding systems within uncertainties. The nuclear modification factors R pPb for both $\overline{\Sigma}^{+}$ and $\overline{\Sigma}^{-}$ are evaluated and compared to those of protons, $Λ$ and $Ξ$ hyperons, and predictions of EPOS LHC and EPOS4 models. No deviations of R pPb for $\overline{\Sigma}^{\pm}$ from the model predictions or measurements for other hadrons are found within uncertainties.

Abualrob, I. J. [University of Houston] (ORCID:000↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Orbitrap LC-MS Analysis of Nanoparticle Composition at the EPCAPE Mount Soledad site between 04 18 2023 and 06 14 2023

Weekly peak lists containing m/z, intensity, and assigned formula for filter samples, size selected for sub-100 nm particles. Filters were collected daily between 4/18/23 and 6/14/23, grouped based on calendar week for extraction, and analyzed via Thermo Scientific Q Exactive Plus Orbitrap LC-MS. Formulas were assigned to background-corrected peak lists and restricted to CHONS/CHONSNa atoms for the negative and positive modes respectively. Filters were grouped into calendar weeks 0-8 with dates provided in README text file.

54 ENVIRONMENTAL SCIENCES↗

CAMFeND: Credibility-Aware Multimodal Fake News Detection with Rotational Attention

In the evolving digital landscape, fake news is a significant challenge, influencing public perception and decision-making. Traditional detection approaches focus on single-modal data or simple multimodal fusion, often overlooking deeper interactions and news credibility. We propose a novel model addressing these limitations by introducing rotational attention and news domain information as a feature. Unlike static attention mechanisms, our rotational attention dynamically shifts query, key, and value roles across text and image inputs, enabling richer cross-modal interaction. Incorporating news domain information further enhances the model’s reliability by associating news posts with top domains extracted from Google search results, reducing false detections. This approach assesses both the content and the broader web context in which the news is discussed. Our model outperforms existing state-of-the-art methods by providing deeper, layered multimodal integration and domain information analysis, resulting in a more robust and adaptive fake news detection system.

Gupta, Nidhi↗

Assurance of Reasoning Enabled Systems (ARES)

ARES was in part motivated by the determination of President’s Council of Advisors on Science and Technology (PCAST) on May 13th, 2023 that published a set of inquiries: In an era in which convincing images, audio, and text can be generated with ease on a massive scale, how can we ensure reliable access to verifiable, trustworthy information? How can we be certain that a particular piece of media is genuinely from the claimed source? What technologies, policies, and infrastructure can be developed to detect and counter AI-generated disinformation? In an effort to automatically analyze and patch/optimize code the work in this report describes various neural Machine Learning (ML) analysis engine implementations to assist in situations where source code is deficient or completely lacking to decompile (lift) binary code to ’C’. The goal is to gradually reduce human intervention. To this end, two Large Language Model (LLM) variants (Code LLama 2, LLama 3.1 and Starcoder1, Starcoder 2) where finetuned with ’before/after’ code pairs on the OpenBLAS library. LLama trained on the lowering process, Starcoder trained on the lifting process with National Security Agency’s (NSA) open-source Ghidra decompiler assist. The inferencing test results indicate correctness for only very short sequences for Starcoder 2. Moving forward, the experiments conclude with a set of recommendations of required resources and technologies

97 MATHEMATICS AND COMPUTING↗

NUM-DAT File Format Specification: Used in M-9 Gun Experiment Data Archiving

The M-9 Shock and Detonation Physics group executes experiments on gun and explosive platforms with large numbers of oscilloscopes used for data acquisition. The data acquisition from these oscilloscopes was automated many years ago using a custom piece of software called RunDig . The default save format from this software is a custom structure referred to as "NUM-DAT" format. This file format includes a text ".DAT" file which is a header file used to interpret the binary ".NUM" file which contains the oscilloscope data. The data save format was originally developed by John Vorthman and has been in use by M-9 personnel for over 20 years. This data format has been used for archiving data from experiments performed by M-9 personnel at TA-40, TA-39, and the TA-55 Impact Test Facility. Numerous custom analysis and visualization programs have also been developed, and continue to be used, that utilize this data format. This document describes the NUM-DAT format and provides code examples for reading the format and converting it to other formats.

47 OTHER INSTRUMENTATION↗

Search for resonances decaying to an anomalous jet and a Higgs boson in proton–proton collisions at $\sqrt{s}=13\,\text {Te}\hspace{-.08em}\text {V}$

This paper presents a search for new physics through the process where a massive particle, X, decays into a Higgs boson and a second particle, Y. The Higgs boson subsequently decays into a bottom quark–antiquark pair, which is reconstructed as a single large-radius jet. The decay products of Yare also assumed to produce a single large-radius jet. The identification of the Yparticle is enhanced by computing the anomaly score of its candidate jet using an autoencoder, which measures deviations from typical quark- or gluon-induced jets. This allows a simultaneous search for multiple Ydecay scenarios within a single analysis. In the main benchmark process, Yis a scalar particle that decays into a Wboson pair. Two other scalar Ydecay processes are also considered as benchmarks: decays to a light quark–antiquark pair, and decays to a top quark–antiquark pair. A fourth benchmark process considers Yas a hadronically decaying top quark, arising from the decay of a vector-like quark into a top quark and a Higgs boson. Data recorded by the CMS experiment at a center-of-mass energy of 13 TeV in 2016–2018, corresponding to an integrated luminosity of 138 fb -1 , are analyzed. The search covers Xmasses between 1.4 and 3.0 TeV and Ymasses between 90 and 400 TeV, with all simulated signals produced in the narrow-width approximation. No significant excess above the standard model background expectation is observed. The most stringent upper limits to date are placed on benchmark signal cross sections for various masses of X and Y particles

Hayrapetyan, A. [Yerevan Physics Institute]↗

HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency

Many workflows in high-energy-physics (HEP) stand to benefit from recent advances in transformer-based large language models (LLMs). While early applications of LLMs focused on text generation and code completion, modern LLMs now support orchestrated agency: the coordinated execution of complex, multi-step tasks through tool use, structured context, and iterative reasoning. We introduce the HEP Toolkit for Agentic Planning, Orchestration, and Deployment (HEPTAPOD), an orchestration framework designed to bring this emerging paradigm to HEP pipelines. The framework enables LLMs to interface with domain-specific tools, construct and manage simulation workflows, and assist in common utility and data analysis tasks through schema-validated operations and run-card-driven configuration. To demonstrate these capabilities, we consider a representative Beyond the Standard Model (BSM) Monte Carlo validation pipeline that spans model generation, event simulation, and downstream analysis within a unified, reproducible workflow. HEPTAPOD provides a structured and auditable layer between human researchers, LLMs, and computational infrastructure, establishing a foundation for transparent, human-in-the-loop systems.

Menzo, Tony [Alabama U.; Fermilab] (ORCID:00000002↗