Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Collinear limit of the four-point energy correlator in N = 4 supersymmetric Yang-Mills theory

We present a compact formula, expressed in terms of classical polylogarithms up to weight three, for the leading order four-point energy correlator in maximally supersymmetric Yang-Mills theory, in the limit where the four detectors are collinear. This formula is derived by combining a simplified, manifestly dual conformal invariant form of the 1 → 4 splitting function obtained from the square of the tree-level five-particle form factor of stress-tensor multiplet operators, with a novel integration-by-parts algorithm operating directly on Feynman parameter integrals. Our results provide valuable data for exploring the structure of physical observables in perturbation theory, and for calculations of jet substructure observables in quantum chromodynamics. Published by the American Physical Society 2024

Chicherin, Dmitry (ORCID:000000028985084X)↗

GENESPACE R Package (GENESPACE) v1.0

In short, the GENESPACE pipeline conducts analysis of orthology networks, constrained within syntenic regions. Since analyses are limited to local tests conducted within syntenic blocks, GENESPACE is agnostic to ploidy, duplicated regions, inversions or other whole-genome chromosomal complexities that are common across many evolutionary lineages. This advantage allows for evolutionary tests in polyploids (e.g. switchgrass, manuscript in review), species with ancient, but retained whole-genome duplications (e.g. pecan, manuscript in prep), high levels of tandem array proliferation (e.g. eukalypts, manuscript in review) and many other factors that can confound comparative genomic analyses. The major advances of GENESPACE are three-fold: First, this is the first R package to integrate visualization and analysis of large-scale comparative genomics. R, which offers a high-level environment for graphical and statistical exploration of data, is often speed- and memory-limited and not used for computationally intensive tasks such as comparative genomics. The highly efficient C++ scripts used in GENESPACE (via data.table) permit a much faster and computationally lightweight implementation of comparative genomics than is currently available. Second, the pipeline itself is novel. To the best of our knowledge, no other program accomplishes synteny-constrained and ploidy-agnostic comparative genomics. Since nearly all plants and many animals have a history of whole-genome duplications, this is a major and necessary advance to the field. Third, GENESPACE offers high-level and intuitive multi-genome graphical outputs. The dotplots and 'riparian' plots produced herein, which are produced entirely through original R code, are publication-ready and easily customizable.

Schmutz, Jeremy↗

Factors Influencing Willingness to Pool in Ride-Hailing Trips

In the past decade, transportation network companies (TNCs) such as Uber, Lyft, and Via have established themselves as a viable transportation alternative to other modes. However, the popularity of these services has come with a fair share of criticism for their negative externalities such as increasing vehicle miles traveled and congestion in cities. Pooled ride-hailing trips, in which all or a part of two individual (or group) trips are combined in and served by a single vehicle, have the potential to reduce these externalities. Pooling of rides is an effective solution to reduce congestion and travel cost, but pooled rides still represent a small percentage of the total trips served (and miles driven) by TNCs relative to single-occupancy (and without customer) vehicle miles. Both TNCs and cities alike will benefit from understanding what factors encourage or deter pooling a ride-hailing trip. In this study, newly available Chicago transportation network provider data were explored to identify the extent to which different socioeconomic, spatiotemporal, and trip characteristics affect willingness to pool (WTP) in ride-hailing trips. Furthermore, multivariate linear regression and machine-learning models were employed to understand and predict WTP based on location, time, and trip factors. The results show intuitive trends, with income level at drop-off and pickup locations and airport trips as the most important predictors of WTP. Results from this study can help TNCs and cities devise strategies that increase pooled ride-hailing, thereby reducing adverse transportation and energy impacts from ride-hailing modes.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Characterization of Pliocene and Miocene Formations in the Wilmington Graben, Offshore Los Angeles, for Large-Scale Geologic Storage of CO2

The project Characterization of Pliocene and Miocene Formations in the Wilmington Graben, Offshore Los Angeles, for Large-Scale Geologic Storage of CO2 is one of 9 site characterization projects that were implemented as part of ARRA (American Recovery and Reinvestment Act). Data from this project was used to improve resolution of data in NATCARB in the area of study. Data related to this study has already been incorporated in NATCARB Atlas. The Los Angeles Basin presents an opportunity for large-scale geologic CO2 storage. Due to its large population and historical and geologic setting as one of the most prolific oil and gas producing basins in the United States, the region is home to more than 12 major power plants and oil refineries that produce more than 5 million metric tons of fossil fuel-related CO2 emissions each year. GeoMechanics Technologies worked to characterize the Pliocene and Miocene sediments in the Wilmington Graben, offshore of Los Angeles, California, for high-volume CO2 storage. The Graben is located offshore of the Los Angeles and Long Beach Harbor area, making it accessible yet geologically isolated from the nearby Wilmington oilfield and onshore areas. These sediments span more than 5,000 feet of vertical interval with an estimated storage resource of more than 100 million metric tons of CO2. The project team analyzed and interpreted existing geologic data within the region, including detailed exploration well log data and 2-D and 3-D seismic data. New seismic lines were acquired to fill in current data gap areas and two new characterization wells were drilled and logged. This information was integrated with existing geologic interpretations for adjacent onshore areas to help characterize optimal areas for CO2 storage and seals to safely store CO2. Integrated 3-D geologic and geomechanical models for the Wilmington Graben were developed to simulate the fate and transport of injected CO2 in the subsurface and to assess risks. This project contributed to the understanding of injectivity, containment mechanisms, rate of dissolution and mineralization, and storage capacity of the Wilmington Graben and associated analogous basins. This effort also provided greater insight into the potential for offshore geologic formations to safely and permanently store CO2.

.las↗

Development of Renewable Energy Data for Facilitating the Clean Energy Transition in Vietnam

Vietnam is transforming its power system with an increasing share of variable renewable energy (VRE) in the generation mix and plans to achieve 27% VRE by capacity in 2030 and 42% by capacity in 2045, as indicated by the draft Eighth Power Development Plan for 2020–-2045 (PDP8). Further increase in VRE penetration is possible with data-driven analysis where availability of data, specifically VRE data, becomes critical. NREL has worked with various stakeholders in Vietnam to create high-resolution multi-year VRE data publicly available to all. This paper reviews this new VRE data for Vietnam and how different stakeholders can use this data to inform VRE deployment decisions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data Fusion to Enhance Quality Control and Analysis with Instruments at the Marine and Coastal Research Laboratory

Deploying environmental monitoring instruments in the marine environment can be challenging, facing challenges around device survivability, biofouling and corrosion, and consistent data collection. This project explores the use of data fusion – the process of integrating multiple data sources to produce more consistent, accurate, and useful information – to build a consistent long-term monitoring system at the Marine and Coastal Research Laboratory (MCRL) in Sequim, Washington. Unused instruments that had been acquired from past projects were inventoried and deployments planned on the MCRL pier and floating dock. A total of 8 instruments were deployed including a tide gauge, hydrophone, acoustic Doppler current profiler (ADCP), photosynthetically active radiation (PAR) sensors, meteorological station, and three water quality sensors. Deployments were planned to be well-protected around the pier structure and a maintenance schedule was created for cleaning and recalibration. An automated data pipeline was created to aggregate data on edge computers that push data to Amazon Web Services (AWS) cloud storage every 15 minutes, performing automated quality control and data transformations using the Time Series Data Analytical Toolkit (TSDAT). Continued efforts are underway to maintain this system into the future, take a data-driven approach to maintenance scheduling, improve the reliability of the system, and share the data with a variety of end-users.

54 ENVIRONMENTAL SCIENCES↗

Combining Spike Time Dependent Plasticity (STDP) and Backpropagation (BP) for Robust and Data Efficient Spiking Neural Networks (SNN)

National security applications require artificial neural networks (ANNs) that consume less power, are fast and dynamic online learners, are fault tolerant, and can learn from unlabeled and imbalanced data. We explore whether two fundamentally different, traditional learning algorithms from artificial intelligence and the biological brain can be merged. We tackle this problem from two directions. First, we start from a theoretical point of view and show that the spike time dependent plasticity (STDP) learning curve observed in biological networks can be derived using the mathematical framework of backpropagation through time. Second, we show that transmission delays, as observed in biological networks, improve the ability of spiking networks to perform classification when trained using a backpropagation of error (BP) method. These results provide evidence that STDP could be compatible with a BP learning rule. Combining these learning algorithms will likely lead to networks more capable of meeting our national security missions.

97 MATHEMATICS AND COMPUTING↗

Floating PV Potential and Technology Validation [Slides]

This presentation covers, at a high-level, the potential for floating solar PV (FPV) in the United States as well as technology validation approaches and research capabilities at the U.S. Department of Energy's National Renewable Energy Laboratory (NREL).

14 SOLAR ENERGY↗

Characterizing the effect of hypersonic boundary layer turbulence on antenna performance: A computational approach

The degradation of antenna performance during hypersonic re-entry is a well known phenomenon that can lead to complete radio blackout. Recent additions to the Empire code establish it as a tool for the study and analysis of the problem. Coupling to the Sandia Parallel Aerodynamics and Reentry Code (SPARC) enables the electromagnetic analysis of realistic re-entry plasma profiles. The geometric flexibility afforded by both Empire and SPARC allow the consideration of arbitrary vehicle and antenna configurations. We have used this tool to study antenna performance during re-entry when the boundary layer becomes turbulent. A concise description of line-of-sight transmissions, which employs advanced statistical methods, was developed. New insights into the low altitude reflectometer readings of RAM-C2 are offered. Techniques for the reconstruction of the re-entry plasma profile from reflectometer data were explored.

42 ENGINEERING↗

Applications of Federated Learning in Semiconductor Manufacturing [Poster]

As semiconductor manufacturers explore advanced data analytics and modeling techniques and data hungry machine learning models increase in popularity due to their accuracy in solving generalized problems and ability to learn complex relationships, federated learning emerges as a privacy preserving machine learning technique for preserving data privacy and ensuring intellectual property protection. Federated Learning is a machine learning technique focused on training models using distributed data that never needs to be centrally stored, allowing the use of advanced machine learning techniques without compromising data privacy, and in the semiconductor manufacturing industry advanced machine learning techniques can reduce cost and time, but maintaining data privacy is essential to maintaining a competitive advantage. This paper systematically reviews existing literature on applications of federated learning in the semiconductor manufacturing industry with a focus on identifying common themes, algorithms, and gaps within the literature to drive future research directions. The findings reveal five key themes, including improvements in quality assurance, virtual models, privacy preservation, reliable data practices, and emerging trends and developments. By identifying key themes in literature on federated learning and semiconductor manufacturing and analyzing gaps and discussed methodologies, this study highlights several potential future research directions to expand the application of federated learning techniques in the semiconductor manufacturing domain.

42 ENGINEERING↗

Form factors for semi-leptonic $B_{(s)} \to D_{(s)}^∗\ell \nu_\ell$ decays

Semileptonic $B_{(s)}$ decays are of great phenomenological interestbecause they allow to extract CKM matrix elements or test lepton flavouruniversality. Taking advantage of existing data, we explore extractingform factors for vector final states using the narrow widthapproximation. Based on RBC/UKQCD's set of 2+1 flavour gauge fieldensembles with Shamir domain-wall fermion and Iwasaki gauge fieldaction, we study semileptonic $B_{(s)}$ decays using domain-wallfermions for light, strange and charm quarks, whereas bottom quarks aresimulated with the relativistic heavy quark (RHQ) action. Exploratoryresults for $B_s \to D_s^* \ell \nu_\ell$ are presented.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A cost and community perspective on the barriers to microbiome data reuse

Microbiome research is becoming a mature field with a wealth of data amassed from diverse ecosystems, yet the ability to fully leverage multi-omics data for reuse remains challenging. To provide a view into researchers’ behavior and attitudes towards data reuse, we surveyed over 700 microbiome researchers to evaluate data sharing and reuse challenges. We found that many researchers are impeded by difficulties with metadata records, challenges with processing and bioinformatics, and problems with data repository submissions. We also explored the cost constraints of data reuse at each step of the data reuse process to better understand “pain points” and to provide a more quantitative perspective from sixteen active researchers. The bioinformatics and data processing step was estimated to be the most time consuming, which aligns with some of the most frequently reported challenges from the community survey. From these two approaches, we present evidence-based recommendations for how to address data sharing and reuse challenges with concrete actions for future work.

59 BASIC BIOLOGICAL SCIENCES↗

A Quick Look at the 3 GHz Radio Sky. II. Hunting for DRAGNs in the VLA Sky Survey

Active galactic nuclei (AGNs) can often be identified in radio images as two lobes, sometimes connected to a core by a radio jet. This multicomponent morphology unfortunately creates difficulties for source finders, leading to components that are (a) separate parts of a wider whole, and (b) offset from the multiwavelength cross identification of the host galaxy. In this work we define an algorithm, DRAGN HUNTER , for identifying double radio sources associated with AGNs (DRAGNs) from component catalog data in the first epoch Quick Look images of the high-resolution (≈3" beam size) Very Large Array Sky Survey (VLASS). We use DRAGN HUNTER to construct a catalog of >17,000 DRAGNs in VLASS for which contamination from spurious sources is estimated at ≈11%. A “high-fidelity” sample consisting of 90% of our catalog is identified for which contamination is <3%. Host galaxies are found for ≈13,000 DRAGNs as well as for an additional 234,000 single-component radio sources. Using these data, we explore the properties of our DRAGNs, finding them to be typically consistent with Fanaroff–Riley class II sources and to allow us to report the discovery of 31 new giant radio galaxies identified using VLASS.

79 ASTRONOMY AND ASTROPHYSICS↗

DoZen Documentation

DoZen (pronounced “do zen”) is for processing, visualizing, and exploring electromagnetic data that is stored in Zonge’s .z3d format. It was created specifically for processing time-lapse controlled source electromagnetic data.

97 MATHEMATICS AND COMPUTING↗

Virtual Volumes

SAND2022-14692 O Virtual Volumes is a CAD2VR plugin that gives users a powerful and intuitive way of exploring volumetric data in a three-dimensional (3D) virtual reality (VR) environment. It enables viewing of and interaction with volumetric and voxel data derived from sources such computed tomography (CT) scans. Current 2D software solutions for viewing CT scans and other volumetric data forces the user to look through "slices" of their data across anatomical planes. In Virtual Volumes, the user can intuitively interact, scale, crop, colorize, and threshold their data in a 3D VR environment. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Krukar, JohnA.↗

Modeling freight mode choice using machine learning classifiers: a comparative study using Commodity Flow Survey (CFS) data

This study explores the usefulness of machine learning classifiers for modeling freight mode choice. We investigate eight commonly used machine learning classifiers, namely Naïve Bayes, Support Vector Machine, Artificial Neural Network, K-Nearest Neighbors, Classification and Regression Tree, Random Forest, Boosting and Bagging, along with the classical Multinomial Logit model. US 2012 Commodity Flow Survey data are used as the primary data source; we augment it with spatial attributes from secondary data sources. The performance of the classifiers is compared based on prediction accuracy results. The current research also examines the role of sample size and training-testing data split ratios on the predictive ability of the various approaches. In addition, the importance of variables is estimated to determine how the variables influence freight mode choice. The results show that the tree-based ensemble classifiers perform the best. Specifically, Random Forest produces the most accurate predictions, closely followed by Boosting and Bagging. With regard to variable importance, shipment characteristics, such as shipment distance, industry classification of the shipper and shipment size, are the most significant factors for freight mode choice decisions.

42 ENGINEERING↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗