Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Applications of Federated Learning in Semiconductor Manufacturing [Poster]

As semiconductor manufacturers explore advanced data analytics and modeling techniques and data hungry machine learning models increase in popularity due to their accuracy in solving generalized problems and ability to learn complex relationships, federated learning emerges as a privacy preserving machine learning technique for preserving data privacy and ensuring intellectual property protection. Federated Learning is a machine learning technique focused on training models using distributed data that never needs to be centrally stored, allowing the use of advanced machine learning techniques without compromising data privacy, and in the semiconductor manufacturing industry advanced machine learning techniques can reduce cost and time, but maintaining data privacy is essential to maintaining a competitive advantage. This paper systematically reviews existing literature on applications of federated learning in the semiconductor manufacturing industry with a focus on identifying common themes, algorithms, and gaps within the literature to drive future research directions. The findings reveal five key themes, including improvements in quality assurance, virtual models, privacy preservation, reliable data practices, and emerging trends and developments. By identifying key themes in literature on federated learning and semiconductor manufacturing and analyzing gaps and discussed methodologies, this study highlights several potential future research directions to expand the application of federated learning techniques in the semiconductor manufacturing domain.

42 ENGINEERING↗

Outage Cause Classification of Power Distribution Systems with Machine Learning and Real-World Data

Power distribution systems are geographically dispersed by nature. It may be affected by various factors, such as vegetation, weather, animal and human behaviors. Present response procedures to an outage event massively rely on expert experience and thus tend to be time-consuming. Automatic outage event detection and classification will help to reduce the responding and restoration time. However, this issue is less addressed with existing research done in this area. In this applied research, a set of waveform pre-processing techniques are first proposed to prepare the waveform data for being used as inputs to the classification algorithm. Further, a machine learning-based algorithm is proposed to classify the outage events according to their root causes, e.g. tree contact, animal contact, lightning, etc. Available data include three phase current & voltage waveforms and contextual information during the distribution system outages. The proposed machine learning algorithm takes the current and voltage waveforms as direct inputs in search of features that humans are unable to capture. Real data provided by a distribution company in the East Tennessee region is used to test the proposed pre-processing techniques and the classification algorithm.

Sun, Haoyuan↗

Radiometric Testing of Germicidal UV Products, Round 2: Upper-Room Luminaires (CALiPER Report)

This report analyzes the independently tested performance of eight germicidal ultraviolet (GUV) upper-room luminaires marketed for use in occupied spaces and purchased between March and June 2023. This type of product is mounted to upper walls or ceilings to treat air in the portion of the room above occupants; this allows for safe use of the room when the device is operating, but requires sufficient air mixing between upper and lower portions of the room. Three of the luminaires used UV-emitting LEDs, and the remaining five luminaires used low-pressure mercury (LPM) lamps. Product testing covered radiometric and electrical performance for each luminaire. Initial performance was measured for all eight products, and four were additionally measured after 100 h and 500 h of operation. Measured performance data allowed for comparison against manufacturer or vendor claims if the tested products included such claims. Some products had no performance data available for a given quantity (e.g., UV-C output power), and only four of the eight luminaires had radiant intensity distribution data files in a standard format (e.g., IES LM-63) available for download from product websites. The lack of publicly available performance data makes it difficult for potential buyers and specifiers to identify suitable products and design GUV systems for their specific applications. When products had performance claims, they were sometimes contradictory (e.g., unexplained differences between multiple power values) or ambiguous (e.g., measurement units conflict with quantity, unclear whether luminaire power or lamp power, unclear whether UV output power or UV-C output power). Three of the eight tested luminaires had claimed output power (i.e., radiant flux) values that exceeded measured values by more than an order of magnitude. There was substantial variation in UV-C radiant efficiency, with a measured range of 0.3–1.9% for LED and 0.4–2.1% for LPM, as shown in Figure 1. For example, the LPM luminaire with 0.4% radiant efficiency would need 5 times the amount of electrical energy used by the LPM luminaire with 2.1% radiant efficiency to produce the same amount of UV-C output power. LPM luminaires that had parabolic reflectors aligned with inclined louvers exhibited substantially higher UV-C radiant efficiency than tested luminaires with other designs, potentially cutting energy use by 75%. These results indicate a substantial opportunity for more energy efficient LPM luminaire designs, while demonstrating that UV LED luminaires can offer comparable UV-C radiant efficiency in this application. This may seem surprising, given that LED emitters have lower UV-C radiant efficiency than LPM lamps, but the efficiency-throttling louvers that are generally required for LPM luminaires typically are not needed for LEDs thanks to their directionality. However, lateral beam angles (which describe beam width as viewed from above) were 41–83° for LED luminaires versus 89–110° for LPM luminaires. More luminaires may be required if their lateral beam angles are relatively small, and coverage may be poor if UV-C radiant intensity distribution (i.e., beam shape) is not considered when designing systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Second-harmonic generation tensors from high-throughput density-functional perturbation theory

Optical materials play a key role in enabling modern optoelectronic technologies in a wide variety of domains such as the medical or the energy sector. Among them, nonlinear optical crystals are of primary importance to achieve a broader range of electromagnetic waves in the devices. However, numerous and contradicting requirements significantly limit the discovery of new potential candidates, which, in turn, hinders the technological development. In the present work, the static nonlinear susceptibility and dielectric tensor are computed via density-functional perturbation theory for a set of 579 inorganic semiconductors. The computational methodology is discussed and the provided database is described with respect to both its data distribution and its format. Several comparisons with both experimental and ab initio results from literature allow to confirm the reliability of our data. The aim of this work is to provide a relevant dataset to foster the identification of promising nonlinear optical crystals in order to motivate their subsequent experimental investigation.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Recent Experience with the CMS Data Management System

The CMS[1] experiment manages a large-scale data infrastructure, currently handling over 200 PB of disk and 500 PB of tape storage and transferring more than 1 PB of data per day on average between various WLCG[2] sites. Utilizing Rucio[3] for high-level data management, FTS[4] for data transfers, and a variety of storage and network technologies at the sites, CMS confronts inevitable challenges due to the system’s growing scale and evolving nature. Key challenges include managing transfer and storage failures, optimizing data distribution across different storages based on production and analysis needs, implementing necessary technology upgrades and migrations, and efficiently handling user requests. The data management team has established comprehensive monitoring to supervise this system and has successfully addressed many of these challenges. The team’s efforts aim to ensure data availability and protection, minimize failures and manual interventions, maximize transfer throughput and resource utilization, and provide reliable user support. This paper details the operational experience of CMS with its data management system in recent years, focusing on the encountered challenges, the effective strategies employed to overcome them and the ongoing challenges as we prepare for future demands.

Öztürk, Hasan [CERN]↗

A High-Throughput Computing Infrastructure to Generate Custom, Open Community Geothermal Datasets

The most significant challenge facing geothermal research, development, and deployment is a lack of comprehensive datasets describing the geological and economical properties of North America. Automated knowledge base construction, the process of designing algorithms to analyze text and images to programmatically build new datasets, is one possible solution to this problem. The xDD library of full-text scientific articles (https://xdd.wisc.edu) is one of the largest collections of open and controlled-access scientific documents available for knowledge base construction in the world, but it has been underutilized by experts in geothermal research. The xDD development team attributed the lack of engagement by software developers and geothermal researchers to two perceived shortcomings of the system. First, the workflow for obtaining data from xDD for local development and testing of data mining applications was unnecessarily abstruse and required significant manual intervention by xDD systems administrators. Second, although xDD already held articles from a broad cross-section of scientific literature with an emphasis on the geosciences, it did not have an explicit set of geothermal research documents that could serve as the nucleus of a geothermal data mining application. To address these issues, the Automated Data Extraction PlaTform (ADEPT) was proposed to extend the data distribution capabilities of the xDD system. The ADEPT extension added the following four key features to xDD: 1) integration of National Geothermal Data System (NGDS) documents into the xDD library to provide an explicitly geothermally-themed collection; 2) improved RESTful (i.e., https-protocol driven) web services for external partners to access xDD data for machine learning application development; 3) a web platform for end-users and xDD administrators to coordinate the development of data mining applications from the initial step of browsing available documents to the final stage of deploying a production-quality machine learning application on high-throughput computing infrastructure; and 4) the development of demonstration data mining applications to illustrate the new workflow to potential collaborators. A total of 21,674 geothermal documents from NGDS were fully ingested into the xDD library and the associated metadata is publicly available through the xDD web services; furthermore, the ADEPT web platform is now publicly accessible and fully live at https://xdd.wisc.edu/adept/.

15 GEOTHERMAL ENERGY↗

A Modern Ethernet Data Acquisition Architecture for Fermilab Beam Instrumentation

The Fermilab Accelerator Division, Instrumentation Department is adopting an open-source framework to replace our embedded VME-based data acquisition systems. Utilizing an iterative methodology, we first moved to embedded Linux, removing the need for VxWorks. Next, we adopted Ethernet on each data acquisition module eliminating the need for the VME backplane in addition to communicating with a rack mount server. Development of DDCP (Distributed Data Communications Protocol), allowed for an abstraction between the firmware and software layers. Each data acquisition module was adapted to read out using 1 GbE and aggregated at a switch which up linked to a 10 GbE network. Current development includes scaling the system to aggregate more modules, to increase bandwidth to support multiple systems and to adopt MicroTCA as a crate technology. The architecture was utilized on various beamlines around the Fermilab complex including PIP2IT, FAST/IOTA and the Muon Delivery Ring. In summary, we were able to develop a data acquisition framework which incrementally replaces VxWorks & VME hardware as well as increases our total bandwidth to 10 Gbit/s using off the shelf Ethernet technology.

43 PARTICLE ACCELERATORS↗

Updates to the n+ 63,65 Cu Evaluations in the Resolved Resonance Region [Slides]

This presentation discusses the motivation and background of the n+ 63,65 Cu Evaluations in the Resolved Resonance Region which is to study the interaction of neutrons with copper as it is important in nuclear applications since critical assembly configurations include metallic copper as reflector. In support to the U.S. Department of Energy (DOE) Nuclear Criticality Safety Program (NCSP), measurements and related evaluations of 63,65 Cu isotopes were selected to improve the agreement with the benchmarks and to assess the importance of the angular distribution data for reactor calculations. Previous and current evaluation work is supported by an experimental campaign initiated before 2010, the 63,65 Cu R-matrix analysis generated resonance parameters up to 300 keV. However, due to outstanding issues in the benchmark performance, ENDF/B-VIII.0 library released a truncated set of resonance parameters up to 100 keV. The goal of this work is to generate an updated set of resonance parameters in the 100-300 keV range to improve the benchmark performance of 63,65 Cu isotopes. In conclusion, R-matrix analysis to update 63,65 Cu evaluations was performed to simultaneously improve benchmark performance and extend the RRR to 300 keV. The benchmark calculations suggest the increased capture cross sections are beneficial, however, further investigation of the measured capture data is needed to understand the large normalization scaling factor needed to improve the reactivity. Also, the copper-reflected benchmarks indicate the need to further investigate angular distributions and extension of RRR to 300 keV is aided well by level statistics considerations. Work to refine the fit of individual resonances is ongoing.

07 ISOTOPE AND RADIATION SOURCES↗

Synthesizing realistic sand assemblies with denoising diffusion in latent space

Abstract The shapes and morphological features of grains in sand assemblies have far‐reaching implications in many engineering applications, such as geotechnical engineering, computer animations, petroleum engineering, and concentrated solar power. Yet, our understanding of the influence of grain geometries on macroscopic response is often only qualitative, due to the limited availability of high‐quality 3D grain geometry data. In this paper, we introduce a denoising diffusion algorithm that uses a set of point clouds collected from the surface of individual sand grains to generate grains in the latent space. By employing a point cloud autoencoder, the three‐dimensional point cloud structures of sand grains are first encoded into a lower‐dimensional latent space. A generative denoising diffusion probabilistic model is trained to produce synthetic sand that maximizes the log‐likelihood of the generated samples belonging to the original data distribution measured by a Kullback‐Leibler divergence. Numerical experiments suggest that the proposed method is capable of generating realistic grains with morphology, shapes and sizes consistent with the training data inferred from an F50 sand database. We then use a rigid contact dynamic simulator to pour the synthetic sand in a confined volume to form granular assemblies in a static equilibrium state with targeted distribution properties. To ensure third‐party validation, 50,000 synthetic sand grains and the 1542 real synchrotron microcomputed tomography (SMT) scans of the F50 sand, as well as the granular assemblies composed of synthetic sand grains are made available in an open‐source repository.

Vlassis, Nikolaos N.↗

Continual Learning for Particle Accelerators

Particle accelerators operate under dynamically changing conditions, which often lead to data distribution drifts. These drifts pose significant challenges for Machine Learning (ML) models, which typically fail to maintain performance when faced with such non-stationary data. In particle accelerators, the primary sources of these data drifts include changes in accelerator settings and non-measured parameters such as machine degradation and environmental factors. Previous research has proposed conditional models to handle multiple beam configurations effectively; however, it is challenging to train the ML models on all possible configuration settings. Additionally, conditional models alone can not address performance degradation caused by drifts due to non-measured factors. These limitations contribute to a significant gap between ML development and its deployment in real-world operational settings. To bridge this gap, in this paper, we identify some of the key areas within particle accelerators where continual learning can help mitigate drift-induced performance degradation. In addition, we present a practical use case where a conditional Auto-Encoder model coupled with memory-based continual learning has been employed to demonstrate stable performance even when underlying data drifts.

Schram, Malachi [Thomas Jefferson National Acceler↗

Continual Learning for Particle Accelerators

Particle accelerators operate under dynamically changing conditions, which often lead to data distribution drifts. These drifts pose significant challenges for Machine Learning (ML) models, which typically fail to maintain performance when faced with such non-stationary data. In particle accelerators, the primary sources of these data drifts include changes in accelerator settings and non-measured parameters such as machine degradation and environmental factors. Previous research has proposed conditional models to handle multiple beam configurations effectively; however, it is challenging to train the ML models on all possible configuration settings. Additionally, conditional models alone can not address performance degradation caused by drifts due to non-measured factors. These limitations contribute to a significant gap between ML development and its deployment in real-world operational settings. To bridge this gap, in this paper, we identify some of the key areas within particle accelerators where continual learning can help mitigate drift-induced performance degradation. In addition, we present a practical use case where a conditional Auto-Encoder model coupled with memory-based continual learning has been employed to demonstrate stable performance even when underlying data drifts

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Operational Evolution of FTS3: A DevOps Driven Approach to Elastic Operations

The File Transfer Service (FTS3) is a distributed data movement service developed at CERN and widely used to transfer data across the Worldwide LHC Computing Grid (WLCG). At Fermilab, FTS3 supports data transfers for multiple experiments, including Intensity Frontier experiments such as DUNE, enabling reliable data movement between WebDAV endpoints in Europe and the Americas.​ At CHEP 2021, we reported on the initial containerized deployment of FTS3 on OKD, the community Kubernetes distribution of Red Hat OpenShift. In this work, we present the subsequent evolution of this deployment, focusing on new operational capabilities introduced to improve scalability, robustness, and long-term maintainability.​ We describe the adoption of more secure and reproducible container build workflows, the integration of DevOps-driven operational practices, and enhancements in monitoring and automation. A key new result is the introduction of horizontal scaling and elastic resource management, allowing FTS3 components to dynamically adapt to workload variations while maintaining service reliability. We also discuss improvements in fault tolerance and operational procedures derived from production experience.​ Finally, we summarize lessons learned from operating FTS3 as a Kubernetes-native service and outline how these developments have improved the resilience and efficiency of data movement operations at Fermilab.

Munoz Flores, Victor Leopoldo [Fermilab]↗

A CIM Based Data Integration Framework for Distribution Utilities

With the proliferation of distributed energy resources and advanced metering, modern electric power distribution systems are data rich and include advanced capabilities for distribution automation. Distribution utilities need applications for planning and operations that can integrate all the available data from the enterprise applications and may incorporate distributed approaches to operate and control. This paper describes a framework for standardizing and integrating the data available at different vendor specific applications in a utility utilizing the Common Information Model (CIM). Theframework incorporates the conversion of these the standardized CIM models into models compatible with GridLAB-D, an open source distribution system simulator, for developing advanced planning and operation strategies.

data integration platform, object-oriented data mo↗

Nonconvex regularization for sparse neural networks

Convex ℓ 1 regularization using an infinite dictionary of neurons has been suggested for constructing neural networks with desired approximation guarantees, but can be affected by an arbitrary amount of over-parametrization. This can lead to a loss of sparsity and result in networks with too many active neurons for the given data, in particular if the number of data samples is large. As a remedy, in this paper, a nonconvex regularization method is investigated in the context of shallow ReLU networks: We prove that in contrast to the convex approach, any resulting (locally optimal) network is finite even in the presence of infinite data (i.e., if the data distribution is known and the limiting case of infinite samples is considered). Moreover, here we show that approximation guarantees and existing bounds on the network size for finite data are maintained.

97 MATHEMATICS AND COMPUTING↗

Distributed Solar 2020 Data Update [Slides]

Berkeley Lab’s Tracking the Sun report summarizes installed prices and other trends among grid-connected, distributed solar photovoltaic (PV) systems in the United States. This report is now being published on a biannual cycle. In 2020, Berkeley Lab has released a more limited Distributed Solar 2020 Data Update, which consists of the same data otherwise published in Tracking the Sun report. The update includes data on more than 1.9 million systems installed through 2019, covering 82% of all distributed PV systems installed nationally through that timeframe.As in prior years, the data update focuses to a large degree on installed prices reported for distributed PV projects, describing both historical trends and variability in pricing across projects.With respect to the historical price trajectory, national median installed prices fell, from 2018 to 2019, by roughly 1% for residential systems, remained essentially flat for small non-residential systems, and fell by 4% for large non-residential systems. Across all three customer segments, these are the slowest annual percentage declines since 2006-2008.Pricing continues to vary widely across individual projects, reflecting, among other things, differences in system sizing and design, installer-level pricing strategies, and local market conditions. For example, among residential systems installed in 2019, the lowest 20% were priced below $3.1/W, while the highest 20% were above $4.5/W. The distributions for non-residential systems exhibit similarly wide spreads.In addition to data on installed prices, the data update also covers a broad range of trends related to distributed PV system design, including: system sizing, module efficiency, module-level power electronics, inverter-loading ratios, solar+storage installations, mounting configuration, panel orientation, third-party ownership, and customer segmentation.

14 SOLAR ENERGY↗

Lowering and Runtime Support for Fortran’s Multi-Image Parallel Features using LLVM Flang, PRIF, and Caffeine

This paper provides an overview of the multi-image parallel features in Fortran 2023 and their implementation in the LLVM flang compiler and the Caffeine parallel runtime library. The features of interest support a Single-Program, Multiple-Data (SPMD) programming model based on executing multiple “images”, each of which is a program instance. The features also support a Partitioned Global Address Space (PGAS) in the form of “coarray” distributed data structures. The paper discusses the lowering of multi-image features to the Parallel Runtime Interface for Fortran (PRIF) and the implementation of PRIF in the Caffeine parallel runtime library. This paper also provides an early view into the design of a new multi-image dialect of the LLVM Multi-Level Intermediate Representation (MLIR). We describe validation and testing of the resulting software stack, and demonstrate that performance compares favorably to another open-source compiler and runtime library: GNU Compiler Collection (GCC) gfortran and OpenCoarrays, respectively.

Bonachea, Dan↗

Computer-aided Abnormality Detection in Chest Radiographs in a Clinical Setting via Domain-adaptation

Deep learning (DL) models are being deployed at medical centers to aid radiologists for diagnosis of lung conditions from chest radiographs. Such models are often trained on a large volume of publicly available labeled radiographs. These pre-trained DL models’ ability to generalize in clinical settings is poor because of the changes in data distributions between publicly available and privately held radiographs. In chest radiographs, the heterogeneity in distributions arises from the diverse conditions in X-ray equipment and their configurations used for generating the images. In the machine learning community, the challenges posed by the heterogeneity in the data generation source is known as domain shift, which is a mode shift in the generative model. In this work, we introduce a domain-shift detection and removal method to overcome this problem. Our experimental results show the proposed method’s effectiveness in deploying a pre-trained DL model for abnormality detection in chest radiographs in a clinical setting.

Dubey, Abhishek↗

Performance Evaluation of Python Based Data Analytics Frameworks in Summit: Early Experiences

The explosion in the volumes of data generated from ever-larger simulation campaigns and experiments or observations necessitates competent tools for data wrangling and analysis). While the Oak Ridge Leadership Computing Facility (OLCF) provides a variety of tools to perform data wrangling and data analysis tasks, Python based tools often lack scalability, or the ability to fully exploit the computational capability of OLCF’s Summit supercomputer. NVIDIA RAPIDS and Dask offer a promising solution to accelerate and distribute data analytics workloads from personal computers to heterogeneous supercomputing systems. We discuss early performance evaluation results of RAPIDS and Dask on Summit to understand their capabilities, scalability, and limitations. Our evaluation includes a subset of RAPIDS libraries, i.e., cuDF, cuML, and cuGraph, and Chainer’s CuPy, and their multi-GPU variants when available.We also draw on the observed trends from the performance evaluation results to discuss best practices for maximizing performance.

Hernandez Arreguin, Benjamin↗