Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Science Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Validation of Modern Nuclear Data Processing in SCALE

The nuclear data (ND) community is continuously developing more accurate, diversified, and comprehensive data for radiation transport modeling to support the nuclear science community. As these community efforts progress, it is crucial that ND processing tools like AMPX (used for SCALE [1] ND) also be developed in parallel to incorporate these new data into transport codes and actually deliver those data to end users. AMPX is a mature, well-tested code that was developed by many people at Oak Ridge National Laboratory (ORNL) over the course of the past few decades. A large portion of the AMPX codebase, however, was outdated, difficult to maintain, and incompatible with modern code development tools. Some of the most important parts of the AMPX code have now been replaced with modern C++ code that can be maintained more cost-effectively and can be tested more rigorously.

AMPX

Global Corn Heat Stress: Mean and SD of Degree Days Above 29°C based on NEX-GDDP-CMIP6 Climate Projections

Description This global dataset provides the estimated mean and standard deviation (SD) of corn heat stress (degree days above 29°C) for a set of climate models in NEX-GDDP-CMIP6 at 0.25-degree resolution. The NEX-GDDP-CMIP6 dataset is comprised of global downscaled climate scenarios derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 6 (CMIP6). The current dataset includes: Long-Term Average Degree Days Above 29°C- Historical Long-Term Average Degree Days Above 29°C- SSP245 Long-Term Standard Deviation of Degree Days Above 29°C- Historical Long-Term Standard Deviation of Degree Days Above 29°C- SSP245 The mean and SD are calculated over 1985-2014 for the historical period and over 2035-2064 for future projections. A full description of methods, including growing season, daily temperature distribution, and statistical coefficients, can be found in Haqiqi (2024). The source climate data are obtained from https://ds.nccs.nasa.gov/thredds2/catalog/catalog.html and are described in Thrasher et al (2022). The codes used to create this dataset are available at https://github.com/ihaqiqi/dd29c_nex_cmip6. Acknowledgments This work was supported by the US Department of Energy, Office of Science, Biological and Environmental Research Program, Earth and Environmental Systems Modeling, MultiSector Dynamics under Cooperative Agreement DE-SC0022141. The data processing, computation, and storage were completed on Purdue Anvil supercomputer and cyberinfrastructure supported by the National Science Foundation HDR award # 2118329: "NSF Institute for Geospatial Understanding through an Integrative Discovery Environment (I-GUIDE)". References Haqiqi. I. (2024). Trade can buffer climate-induced risks and volatilities in crop supply. Environmental Research: Food Systems. https://doi.org/10.1088/2976-601X/ad7d12 Thrasher, B., Wang, W., Michaelis, A., Melton, F., Lee, T., & Nemani, R. (2022). NASA global daily downscaled projections, CMIP6. Scientific Data, 9(1), 262. https://doi.org/10.1038/s41597-022-01393-4

Climate Change

RTN-117: Image Calibration and Instrument Signal Removal for the First Year of the LSST

The NSF-DOE Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) requires calibration products that provide uniform, stable, and accurate photometric and astrometric performance across the 3.2 gigapixel focal plane and throughout the 10-year survey. This paper details the algorithms and workflows used to produce instrument calibrations and remove instrumental artifacts for Data Preview 2 (DP2)---the first end-to-end processing demonstration using on-sky data with the LSST Camera (LSSTCam). We describe the verification, acceptance, and certification framework used to assess calibration quality and quantify residual systematics. We show baseline metrics on calibrated science images to evaluate the robustness of the calibration and instrument signature removal (ISR) data processing pipelines for DP2. Finally, we summarize the known limitations observed in DP2 production and outline expected algorithmic improvements for the first public LSST data release (Data Release~1, DR1).

79 ASTRONOMY AND ASTROPHYSICS

NANT Site - ASSIST / Processed Data

This dataset contains processed summary data from the ASSIST operated by NOAA Physical Science Laboratory on Nantucket Island for WFIP3.

17 WIND ENERGY

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes

“One Table to Rule Them All”: How a Single Table can Enable Extensive Insights, Analytics and Assessment on Human Mobility Data

While much research has been conducted in Human Mobility Science, most studies on the analytics/insights part generally focus on one of the following: processing and analytics on human stop-trip behavior, design of individual mobility metrics (often in silos), calculation and characterization of only a handful (typically 5-6) of human mobility metrics on geospatial-temporal human mobility data of interest. Although human mobility research offers a vast and diverse array of available metrics, most individual studies typically compute only a small subset of five or six metrics at a time when analyzing trajectory datasets of human mobility across different areas of interest. This paper is motivated by the critical need to repeatedly compute an extensive array of human mobility metrics across several trajectory datasets and perform individual metric-level benchmarking to establish a new, standardized Test and Evaluation (T&E) suite for the field of Human Mobility Science. We first present our findings on the minimal yet sufficient pre-processing required to reliably and efficiently compute a wide range of human mobility metrics. The key findings are specifically related to the proposed Composite Stop Locations table, which serves as a core pre-processing data layer. Subsequently, we present a case study demonstrating how the Composite Stop Locations table facilitates computation of at least 14 distinct human mobility metrics (unlike 5-6 different set of metrics used for studies in the literature) using the popular and open-source OpenPFLOW dataset. Finally, we have also presented an example of our benchmarking methodology to evaluate the quality and performance of the trajectory dataset of interest, assessed across multiple human mobility metrics.

De, Debraj [ORNL] (ORCID:0000000233630020)

Advanced Interactive 3D Visualization Tool for Customizable Analyses of Tomography Datasets in Material Science

Current methods for visualizing and analyzing 3D tomography datasets in materials science often lack the interactivity and depth required for detailed structural insights. This limitation restricts a researchers' ability to accurately interpret complex data, which is critical for advancing material innovations and understanding structural properties. To address this issue, we have developed a novel, web-based interactive 3D visualization and analysis tool from the Trame framework that offers customizable features to enhance data interpretability. The tool allows users to adjust parameters such as visible range, slice planes, data rotation, and layering, providing a more detailed and dynamic view of complex structures. Its user-friendly web interface increases the accessibility and ease of use for both novice and experienced researchers, to visualize large volumetric datasets. The tool supports a diverse range of data formats, making it versatile for various research applications. Unique capabilities include real-time data manipulation, automated feature detection, context-sensitive feedback, and real-time volume calculations and distributions per sliced region or layer, alongside the ability to quickly generate high-quality screenshots and videos for presentations and reports. These advancements offer a comprehensive solution for enhanced 3D data exploration, significantly improving the analysis process and communication of results in materials science.

36 - MATERIALS SCIENCE

Recent progress in atomic-scale controlled plasma processing

Atomic-scale control in plasma processing is becoming increasingly critical for fabricating of advanced semiconductor devices, particularly as the industry shifts toward three-dimensional (3D) architectures and high-aspect-ratio (HAR) structures. This review presents a comprehensive overview of recent developments in atomic-scale controlled plasma processes, organized along two key directions: the hierarchical structure of plasma–surface interactions and the generational evolution of atomic layer processing (ALP) technologies. We examined the gas phase, where molecular design enables selective generation of ions and radicals; the boundary layer, where transport phenomena govern species delivery into nanoscale features, and the surface, where temperature-dependent reactions and cyclic processing determine etching selectivity and precision. Building on this foundation, we outline five generations of ALP—from thermal atomic layer deposition to transport-aware, temporally and structurally decoupled processes—highlighting the increasing sophistication of process control. The review further explores the transition from empirical recipe development to science-based, data-driven methodologies. By integrating quantum-chemical modeling, advanced diagnostics, and machine learning, we demonstrated how predictive models can link plasma species composition to process outcomes, enabling autonomous and adaptive control strategies. Finally, this review discusses the broader societal implications of plasma process innovation through the E4 quartet: energy and resource efficiency, environmental sustainability, evolutionary advancement, and educational promotion. These principles guide the development of sustainable and intelligent atomic-scale manufacturing technologies that are not only technically advanced but also socially responsible.

Ishikawa, Kenji [Nagoya Univ. (Japan)] (ORCID:0000

Real-time data processing for serial crystallography experiments

We report the use of streaming data interfaces to perform fully online data processing for serial crystallography experiments, without storing intermediate data on disk. The system produces Bragg reflection intensity measurements suitable for scaling and merging, with a latency of less than 1 s per frame. Our system uses the CrystFEL software in combination with the ASAP::O data framework. In a series of user experiments at PETRA III, frames from a 16 megapixel Dectris EIGER2 X detector were searched for peaks, indexed and integrated at the maximum full-frame readout speed of 133 frames per second. The computational resources required depend on various factors, most significantly the fraction of non-blank frames ('hits'). The average single-thread processing time per frame was 242 ms for blank frames and 455 ms for hits, meaning that a single 96-core computing node was sufficient to keep up with the data, with ample headroom for unexpected throughput reductions. Further significant improvements are expected, for example by binning pixel intensities together to reduce the pixel count. We discuss the implications of real-time data processing on the `data deluge' problem from recent and future photon-science experiments, in particular on calibration requirements, computing access patterns and the need for the preservation of raw data.

47 OTHER INSTRUMENTATION

Advances in the photon avalanche luminescence of inorganic lanthanide-doped nanomaterials

Photon avalanche (PA)—where the absorption of a single photon initiates a ‘chain reaction’ of additional absorption and energy transfer events within a material—is a highly nonlinear optical process that results in upconverted light emission with an exceptionally steep dependence on the illumination intensity. Over 40 years following the first demonstration of photon avalanche emission in lanthanide-doped bulk crystals, PA emission has been achieved in nanometer-scale colloidal particles. The scaling of PA to nanomaterials has resulted in significant and rapid advances, such as luminescence imaging beyond the diffraction limit of light, optical thermometry and force sensing with (sub)micron spatial resolution, and all-optical data storage and processing. In this review, we discuss the fundamental principles underpinning PA and survey the studies leading to the development of nanoscale PA. Finally, we offer a perspective on how this knowledge can be used for the development of next-generation PA nanomaterials optimized for a broad range of applications, including mid-IR imaging, luminescence thermometry, (bio)sensing, optical data processing and nanophotonics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

BNF Radar b1 Data Processing Report: Spring 2025

The U.S. Department of Energy’s Atmospheric Radiation Measurement (ARM) User Facility supports atmospheric science through an integrated network of fixed and mobile observatories. These facilities collect continuous and campaign-based observations of atmospheric properties, with the goal of improving the representation of clouds, aerosols, precipitation, and radiation in Earth system models. The Bankhead National Forest (BNF) site, established as an ARM Mobile Facility (AMF) on 1 October 2024, is situated in a forested region of northern Alabama. Its strategic location in a southeastern U.S. environment characterized by complex terrain, diverse land cover, and frequent convective storms provides a valuable opportunity to examine coupled land-atmosphere processes under natural variability.

54 ENVIRONMENTAL SCIENCES

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence

Addressing Limitations of the Endpoint Slippage Analysis

Some rate of oxidation and reduction side-reactions will inevitably coexist in most rechargeable batteries. While parasitic reduction traps electrons, parasitic oxidation donates electrons to the cell’s inventory and may cause temporary capacity gain. Consequently, capacity measurements can provide unreliable information about the total extent of side-reactions occurring in the cell. The most widely used method to determine the rate of both these parasitic processes involves analyzing the slippage of endpoints, which consists in tracking the termination of cell charge and discharge when data is represented along a cumulative capacity axis. Here, we argue that this approach could lead to inaccuracies when applied to certain systems, which includes Si electrodes in Li-ion batteries and hard carbon in Na-ion batteries. This inaccuracy originates from the smooth nature of the voltage profiles of these materials at low and high alkali-ion content, causing the termination of charge and discharge to be dictated by voltage changes at both the positive and negative electrodes. We analyze this issue in quantitative terms and propose equations that can provide true rates of parasitic processes from experimental endpoint slippage data. This work shows that, in battery science, well-established analytical approaches may not be directly transferrable to new electrode systems.

25 ENERGY STORAGE

The DECam MAGIC Survey $-$ Mapping the Ancient Galaxy in CaHK: Overview and Summary of Early Science

We present the DECam Mapping the Ancient Galaxy in CaHK (MAGIC) survey, a 54-night NOIRLab Survey Program to image $\gtrsim$5,000$\,$deg$^2$ of the southern hemisphere using a metallicity-sensitive narrow-band filter covering the Ca$\,$ii$\,$H&K lines centered at 3955$\,$A. This filter is installed on the Dark Energy Camera (DECam), mounted on the 4-m NSF Víctor M. Blanco Telescope. The survey reaches typical $10σ$ depths of $\text{mag}_{\text{CaHK}} \approx 22.5$, 3$-$4$\,$mag deeper than comparable surveys in the southern hemisphere. By combining photometry from this Ca$\,$ii$\,$H&K filter with existing DECam $g,r,i$ broadband photometry from the DECam Local Volume Exploration (DELVE) survey, MAGIC is deriving photometric metallicities for red giant branch stars down to the magnitude limit of usable proper motions from Gaia data release 3 (DR3). MAGIC has already imaged $\sim$3,000$\,$deg$^2$, supplemented by other affiliated observing programs that have used this filter to image star clusters, dwarf galaxies, and stellar streams. We overview MAGIC's survey strategy, describe data processing through the derivation of metallicities and photometric distances, and summarize early science results that have been published with this dataset. In addition, we present several new results, including the confirmation of a distant ($>5\,r_h$) member of the Reticulum II ultra-faint dwarf galaxy, on-sky density maps of low-metallicity stars into the distant Milky Way halo ($\sim150\,$kpc) recovering 13/14 ultra-faint dwarf galaxies in the current footprint, and a validation of our initial targeting of extremely metal-poor stars. Collectively, these results demonstrate that the MAGIC dataset enables cutting-edge studies of the faint, low-metallicity regime of the Milky Way and its substructures.

Chiti, A. [KIPAC, Menlo Park] (ORCID:0000000271556

New York-Presbyterian and Columbia University Irving Medical Center Requirements (Analysis Report)

EPOC uses the Deep Dive process to discuss and analyze current and planned science, research, or education activities and the anticipated data output of a particular use case, site, or project to help inform the strategic planning of a campus or regional networking environment. This includes understanding future needs related to network operations, network capacity upgrades, and other technological service investments. A Deep Dive comprehensively surveys major research stakeholders’ plans and processes in order to investigate data management requirements over the next 5–10 years. Between February and June 2024, staff members from the Engagement and Performance Operations Center (EPOC) met with researchers and staff from New York-Presbyterian (NYP), Columbia University Irving Medical Center (CUIMC), and NYSERNet for the purpose of a Deep Dive into scientific and research drivers. The goal of this activity was to help characterize the requirements for a number of campus use cases, and to enable cyberinfrastructure support staff to better understand the needs of the researchers within the community. Material for this event included the written documentation from each of the profiled research areas, documentation about the current state of technology support, and a write-up of the discussion that took place via e-mail and video conferencing. The case studies highlighted the ongoing challenges and opportunities that NYP and CUIMC have in supporting a cross-section of established and emerging research use cases. Each case study mentioned unique challenges which were summarized into common needs.

97 MATHEMATICS AND COMPUTING