Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “task tuning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Structured information extraction from scientific text with large language models

Extracting structured knowledge from scientific text remains a challenging task for machine learning models. Here, we present a simple approach to joint named entity recognition and relation extraction and demonstrate how pretrained large language models (GPT-3, Llama-2) can be fine-tuned to extract useful records of complex scientific knowledge. We test three representative tasks in materials chemistry: linking dopants and host materials, cataloging metal-organic frameworks, and general composition/phase/morphology/application information extraction. Records are extracted from single sentences or entire paragraphs, and the output can be returned as simple English sentences or a more structured format such as a list of JSON objects. This approach represents a simple, accessible, and highly flexible route to obtaining large databases of structured specialized scientific knowledge extracted from research papers.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Performance Profile of Transformer Fine-Tuning in Multi-GPU Cloud Environments

The study presented here focuses on performance characteristics and trade-offs associated with running machine-learning tasks in multi-GPU environments on both on-site cloud computing resources and commercial cloud services (Azure). Specifically, this study examines these tradeoffs by examining the performance of training and fine-tuning of transformer-based deep-learning (DL) networks on clinical notes and data, a task of critical importance in the medical domain. To this end, we perform DL-related experiments on the widely deployed NVIDIA V100 GPUs and on the newer A100 GPUs connected via NVLink or PCIe. This study analyzes the execution time of major operations to train DL models and investigate popular options to optimize each of them. We examine and present the findings on the impacts that various operations (e.g. data loading into GPUs, training, fine-tuning), optimizations, and system configurations (single vs. multi-GPU, NVLink vs. PCIe) have on the overall training performance.

Begoli, Edmon↗

Neural architecture codesign for fast physics applications

We develop a pipeline to streamline neural architecture codesign for physics applications to reduce the need for ML expertise when designing models for novel tasks. Our method employs neural architecture search and network compression in a two-stage approach to discover hardware efficient models. This approach consists of a global search stage that explores a wide range of architectures while considering hardware constraints, followed by a local search stage that fine-tunes and compresses the most promising candidates. We exceed performance on various tasks and show further speedup through model compression techniques such as quantization-aware-training and neural network pruning. We synthesize the optimal models to high level synthesis code for FPGA deployment with the hls4ml library. Additionally, our hierarchical search space provides greater flexibility in optimization, which can easily extend to other tasks and domains. We demonstrate this with two case studies: Bragg peak finding in materials science and jet classification in high energy physics, achieving models with improved accuracy, smaller latencies, or reduced resource utilization relative to the baseline models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Study of Overfitting by Machine Learning Methods Using Generalization Equations

The training error of Machine Learning (ML) methods has been extensively used for performance assessment, and its low values have been used as a main justification for complex methods such as estimator fusion and ensembles, and hyper parameter tuning. We present two practical cases where independent tests indicate that the low training error is more of a reflection of over-fitting rather than the generalization ability. We derive a generic form of the generalization equations that separates the training error terms of ML methods from their epistemic terms that correspond to approximation and learnability properties. It provides a framework to separately account for both terms to ensure an overall high generalization performance. For regression estimation tasks, we derive conditions for performance enhancements achieved by hyper parameter tuning, and fusion and ensemble methods over their constituent methods. We present experimental measurements and ML estimates that illustrate the analytical results for the throughput profile estimation of a data transport infrastructure.

Rao, Nageswara↗

Online transfer learning strategy for enhancing the scalability and deployment of deep reinforcement learning control in smart buildings

In recent years, advanced control strategies based on Deep Reinforcement Learning (DRL) proved to be effective in optimizing the management of integrated energy systems in buildings, reducing energy costs and improving indoor comfort conditions when compared to traditional reactive controllers. However, the scalability and implementation of DRL controllers are still limited since they require a considerable amount of time before converging to a near-optimal solution. This issue is currently addressed in literature through the offline pre-training of the DRL agent. However this solution results in two main critical issues: (1) the need to develop a building surrogate model to perform the training task, and (2) the need to perform a fine-tuning process over several training episodes to obtain a near-optimal control policy. In this context, this paper introduces an Online Transfer Learning (OTL) strategy that exploits two knowledge-sharing techniques, weight-initialization and imitation learning, to transfer a DRL control policy from a source office building to various target buildings in a simulation environment coupling EnergyPlus and Python. A DRL controller based on discrete Soft Actor–Critic (SAC) is trained on the source building to manage the operation of a cooling system consisting of a chiller and a thermal storage. Several target buildings are defined to benchmark the performance of the OTL strategy with that of a Rule-Based Controller (RBC) and two DRL-based control strategies, deployed in offline and online fashion. The strategy adopted for OTL emulates the real world implementation with a simulation process by implementing the transferred DRL agent for a single episode in the target buildings. Target buildings have the same geometrical features and are served by the same energy system as the source building, but differ in terms of weather conditions, electricity price schedules, occupancy patterns, and building envelope efficiency levels. The results show that the OTL strategy can reduce the cumulated sum of temperature violations on average by 50% and 80% respectively when compared to RBC and online DRL while enhancing the energy system operation with electricity cost savings ranging between 20% and 40%. Furthermore, the OTL agent performs slightly worse than the offline DRL controller but it does not require any modeling effort and can be implemented directly on target buildings emulating a real-world implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

De novo design of protein structure and function with RFdiffusion

Abstract There has been considerable recent progress in designing new proteins using deep-learning methods 1–9 . Despite this progress, a general deep-learning framework for protein design that enables solution of a wide range of design challenges, including de novo binder design and design of higher-order symmetric architectures, has yet to be described. Diffusion models 10,11 have had considerable success in image and language generative modelling but limited success when applied to protein modelling, probably due to the complexity of protein backbone geometry and sequence–structure relationships. Here we show that by fine-tuning the RoseTTAFold structure prediction network on protein structure denoising tasks, we obtain a generative model of protein backbones that achieves outstanding performance on unconditional and topology-constrained protein monomer design, protein binder design, symmetric oligomer design, enzyme active site scaffolding and symmetric motif scaffolding for therapeutic and metal-binding protein design. We demonstrate the power and generality of the method, called RoseTTAFold diffusion (RFdiffusion), by experimentally characterizing the structures and functions of hundreds of designed symmetric assemblies, metal-binding proteins and protein binders. The accuracy of RFdiffusion is confirmed by the cryogenic electron microscopy structure of a designed binder in complex with influenza haemagglutinin that is nearly identical to the design model. In a manner analogous to networks that produce images from user-specified inputs, RFdiffusion enables the design of diverse functional proteins from simple molecular specifications.

Science & Technology - Other Topics↗

Automatic Qubit Characterization and Gate Optimization with QubiC

As the size and complexity of a quantum computer increases, quantum bit (qubit) characterization and gate optimization become complex and time-consuming tasks. Current calibration techniques require complicated and verbose measurements to tune up qubits and gates, which cannot easily expand to the large-scale quantum systems. We develop a concise and automatic calibration protocol to characterize qubits and optimize gates using QubiC, which is an open source FPGA (field-programmable gate array) based control and measurement system for superconducting quantum information processors. We propose multi-dimensional loss-based optimization of single-qubit gates and full XY-plane measurement method for the two-qubit CNOT gate calibration. We demonstrate the QubiC automatic calibration protocols are capable of delivering high-fidelity gates on the state-of-the-art transmon-type processor operating at the Advanced Quantum Testbed at Lawrence Berkeley National Laboratory. Finally, the single-qubit and two-qubit Clifford gate infidelities measured by randomized benchmarking are of 4.9(1.1) × 10 -4 and 1.4(3) × 10 -2 , respectively.

97 MATHEMATICS AND COMPUTING↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Using Crystallization to Control Filler Dispersion and Vice Versa in Polymer Nanocomposites

Our overall goal was to exploit our ability to control the spatial dispersion of spherical nanoparticles in a semicrystalline polymer matrix to tune the properties of the resulting nanocomposite. In particular, our two tasks were: Hierarchically organize spherical nanoparticles into unique morphologies by varying the ratio of the polymer crystallization rate to the filler diffusion rate during matrix solidification. Organize nanoparticles in the matrix melt into superstructrues as a means to tailor polymer crystallization rates, and potentially the crystal structure and/or morphology of the semicrystalline matrix.

36 MATERIALS SCIENCE↗

Fast Local Spatial Verification for Feature-Agnostic Large-Scale Image Retrieval

Images from social media can reflect diverse viewpoints, heated arguments, and expressions of creativity, adding new complexity to retrieval tasks. Researchers working on Content-Based Image Retrieval (CBIR) have traditionally tuned their algorithms to match filtered results with user search intent. However, we are now bombarded with composite images of unknown origin, authenticity, and even meaning. With such uncertainty, users may not have an initial idea of what the search query results should look like. For instance, hidden people, spliced objects, and subtly altered scenes can be difficult for a user to detect initially in a meme image, but may contribute significantly to its composition. It is pertinent to design systems that retrieve images with these nuanced relationships in addition to providing more traditional results, such as duplicates and near-duplicates — and to do so with enough efficiency at large scale. In this work, we propose a new approach for spatial verification that aims at modeling object-level regions using image keypoints retrieved from an image index, which is then used to accurately weight small contributing objects within the results, without the need for costly object detection steps. We call this method the Objects in Scene to Objects in Scene (OS2OS) score, and it is optimized for fast matrix operations, which can run quickly on either CPUs or GPUs. It performs comparably to state-of-the-art methods on classic CBIR problems (Oxford 5K, Paris 6K, and Google-Landmarks), and outperforms them in emerging retrieval tasks such as image composite matching in the NIST MFC2018 dataset and meme-style imagery from Reddit.

42 ENGINEERING↗

High precision strike point control to support experiments in the DIII-D small angle slot divertor

Improvements in strike point control enabled tests with the new Small Angle Slot (SAS) divertor installed at DIII-D with precision as tight as ±1 mm, barring brief perturbations from ELMs. Aside from being narrower than previous baffle structures in DIII-D, simulations indicate that performance of the SAS divertor should be very sensitive to strike point position, making the improved strike point control essential for experiments with the SAS divertor. Improved control entails regulating more divertor geometry target locations than before and setting higher gain on the strike point target. This is achieved by tasking more poloidal field (PF) shaping coils with divertor topology control, fine tuning of the coil configuration and algorithm gains, and adding regulation of the power supply common rail. Here, the result is that although more variables are controlled, precision is better: high frequency jitter in the strike point position is lower on average compared to related discharges with standard divertor shape control.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predicting seismic amplitudes with machine learning

The accurate estimation of seismic wave amplitude is vital to precisely determine the yield, magnitude, and event discrimination possible for a given network – a critical element in nuclear explosion monitoring. This task is complicated by several factors, including but not limited to radiation pattern, scattering effects, and crustal variations, which can lead to the attenuation or amplification of amplitude along a given raypath. In this report, we explore the novel application of machine learning to the task of seismic amplitude estimation by training a simple Artificial Neural Network (ANN) on an S-wave amplitude dataset from Lai et al. (2019). Attributes from this dataset used as input to the ANN included event-station distances, station locations (latitude, longitude), event locations (latitude, longitude), event depths, event magnitudes, radiation patterns, signal-to noise ratio (SNR) measurements (average-amplitude, peak-to-trough, maximum peak), and signal periods. We find that the trained ANN predicts S-wave amplitudes with a modest tendency toward underestimating the actual values, as indicated by a linear regression between predicted and actual data (slope: 0.892, intercept: -0.651). These results suggest that an ANN can perform this task, with potential for significant improvements through improved datasets, architectures, and parameter tuning.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Microscopic Imprints of Learned Solutions in Tunable Networks

In physical networks trained using supervised learning, physical parameters are adjusted to produce desired responses to inputs. An example is an electrical contrastive local learning network of nodes connected by edges that adjust their conductances during training. When an edge conductance changes, it upsets the current balance of every node. In response, physics adjusts the node voltages to minimize the dissipated power. Learning in these systems is therefore a coupled double-optimization process, in which the network descends both a cost landscape in the high-dimensional space of edge conductances and a physical landscape—the power dissipation—in the high-dimensional space of node voltages. Because of this coupling, the physical landscape of a trained network contains information about the learned task. Here, we derive a structure-function relation for trained tunable networks and demonstrate that all the physical information relevant to the trained input-output relation can be captured by a tuning susceptibility, an experimentally measurable quantity. We supplement our theoretical results with simulations to show that the tuning susceptibility is correlated with functional importance and that we can extract physical insight into how the system performs the task from the conductances of highly susceptible edges. Our analysis is general and can be applied directly to mechanical networks, such as networks trained for protein-inspired function such as allostery.

36 MATERIALS SCIENCE↗

A Software Framework for Comparing Training Approaches for Spiking Neuromorphic Systems

There are a wide variety of training approaches for spiking neural networks for neuromorphic deployment. However, it is often not clear how these training algorithms perform or compare when applied across multiple neuromorphic hardware platforms and multiple datasets. In this work, we present a software framework for comparing performance across four neuromorphic training algorithms across three neuromorphic simulators and four simple classification tasks. We introduce an approach for training a spiking neural network using a decision tree, and we compare this approach to training algorithms based on evolutionary algorithms, back-propagation, and reservoir computing. We present a hyperparameter optimization approach to tune the hyperparameters of the algorithm, and show that these optimized hyperparameters depend on the processor, algorithm, and classification task. Finally, we compare the performance of the optimized algorithms across multiple metrics, including accuracy, training time, and resulting network size, and we show that there is not one best training algorithm across all datasets and performance metrics.

Schuman, Catherine↗

Kinetic Growth of Multicomponent Microcompartment Shells

An important goal of systems and synthetic biology is to produce high value chemical species in large quantities. Microcompartments, which are protein nanoshells encapsulating catalytic enzyme cargo, could potentially function as tunable nanobioreactors inside and outside cells to generate these high value species. Modifying the morphology of microcompartments through genetic engineering of shell proteins is one viable strategy to tune cofactor and metabolite access to encapsulated enzymes. However, this is a difficult task without understanding how changing interactions between the many different types of shell proteins and enzymes affect microcompartment assembly and shape. Here, we use multiscale molecular dynamics and experimental data to describe assembly pathways available to microcompartments composed of multiple types of shell proteins with varied interactions. As the average interaction between the enzyme cargo and the multiple types of shell proteins is weakened, the shell assembly pathway transitions from (i) nucleating on the enzyme cargo to (ii) nucleating in the bulk and then binding the cargo as it grows to (iii) an empty shell. Atomistic simulations and experiments using the 1,2-propanediol utilization microcompartment system demonstrate that shell protein interactions are highly varied and consistent with our multicomponent, coarse-grained model. Furthermore, our results suggest that intrinsic bending angles control the size of these microcompartments. Altogether, our simulations and experiments provide guidance to control microcomparmtent size and assembly by modulating the interactions between shell proteins.

defects↗

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

Automatic Tuner for the Step Sizes of Gradient-Based RT-OPF DERMS Control Systems [SWR-22-50]

Gradient-based controllers have parameters called step sizes that determine how sensitive its control actions are to received inputs. Finding a practical setting for a step size is important. If it is too small, the controller is ineffective by not reacting in a significant manor to its received inputs. On the other hand, if it is too large, the actions implemented by the controller may be excessive and cause the controlled system to become unstable. The main objective of the Automatic Tuner for the Step Sizes of Gradient-Based RT-OPF DERMS Control Systems is to speed up the implementation of gradient-based control schemes by automatically tuning its step sizes instead of having a practitioner spend considerable time doing the task manually. Only simple adjustments need to be made to the code of each controller with a step size to allow the addition of an individual automatic tuner. This can be applied to any RT-OPF based control for DER management.

Comden, Joshua↗

A kinetic-based regularization method for data science applications

We propose a physics-based regularization technique for function learning, inspired by statistical mechanics. By drawing an analogy between optimizing the parameters of an interpolator and minimizing the energy of a system, we introduce corrections that impose constraints on the lower-order moments of the data distribution. This minimizes the discrepancy between the discrete and continuum representations of the data, in turn allowing to access more favorable energy landscapes, thus improving the accuracy of the interpolator. Our approach improves performance in both interpolation and regression tasks, even in high-dimensional spaces. Unlike traditional methods, it does not require empirical parameter tuning, making it particularly effective for handling noisy data. We also show that thanks to its local nature, the method offers computational and memory efficiency advantages over Radial Basis Function interpolators, especially for large datasets.

97 MATHEMATICS AND COMPUTING↗