Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for scientific computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

SatNet: A Benchmark for Satellite Scheduling Optimization

Satellites provide essential services such as networking and weather tracking, and the number of near-earth and deep space satellites are expected to grow rapidly in the coming years. Communications with terrestrial ground stations is one of the critical functionalities of any space mission. Satellite scheduling is a problem that has been scientifically investigated since the 1970s. A central aspect of this problem is the need to consider resource contention and satellite visibility constraints as they require line of sight. Due to the combinatorial nature of the problem, prior solutions such as linear programs and evolutionary algorithms require extensive compute capabilities to output a feasible schedule for each scenario. Machine learning based scheduling can provide an alternative solution by training a model with historical data and generating a schedule quickly with model inference. We present SatNet, a benchmark for satellite scheduling optimization based on historical data from the NASA Deep Space Network. We propose formulation of the satellite scheduling problem as a Markov Decision Process and use reinforcement learning (RL) policies to generate schedules. The nature of constraints imposed by SatNet differ from other combinatorial optimization problems such as vehicle routing studied in prior literature. Our initial results indicate that RL is an alternative optimization approach that can generate candidate solutions of comparable quality to existing state-of-the-practice results. However, we also find that RL policies overfit to the training dataset and do not generalize well to new data, thereby necessitating continued research on reusable and generalizable agents.

Wilson, Brian↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

ExaWorks software development kit: a robust and scalable collection of interoperable workflows technologies

Scientific discovery increasingly requires executing heterogeneous scientific workflows on high-performance computing (HPC) platforms. Heterogeneous workflows contain different types of tasks (e.g., simulation, analysis, and learning) that need to be mapped, scheduled, and launched on different computing. That requires a software stack that enables users to code their workflows and automate resource management and workflow execution. Currently, there are many workflow technologies with diverse levels of robustness and capabilities, and users face difficult choices of software that can effectively and efficiently support their use cases on HPC machines, especially when considering the latest exascale platforms. We contributed to addressing this issue by developing the ExaWorks Software Development Kit (SDK). The SDK is a curated collection of workflow technologies engineered following current best practices and specifically designed to work on HPC platforms. We present our experience with (1) curating those technologies, (2) integrating them to provide users with new capabilities, (3) developing a continuous integration platform to test the SDK on DOE HPC platforms, (4) designing a dashboard to publish the results of those tests, and (5) devising an innovative documentation platform to help users to use those technologies. Our experience details the requirements and the best practices needed to curate workflow technologies, and it also serves as a blueprint for the capabilities and services that DOE will have to offer to support a variety of scientific heterogeneous workflows on the newly available exascale HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Understanding and Leveraging the I/O Patterns of Emerging Machine Learning Analytics

The scientific community is currently experiencing unprecedented amounts of data generated by cutting-edge science facilities. Soon facilities will be producing up to 1 PB/s which will force scientist to use more autonomous techniques to learn from the data. The adoption of machine learning methods, like deep learning techniques, in large-scale workflows comes with a shift in the workflow’s computational and I/O patterns. These changes often include iterative processes and model architecture searches, in which datasets are analyzed multiple times in different formats with different model configurations in order to find accurate, reliable and efficient learning models. This shift in behavior brings changes in I/O patterns at the application level as well at the system level. These changes also bring new challenges for the HPC I/O teams, since these patterns contain more complex I/O workloads. In this paper we discuss the I/O patterns experienced by emerging analytical codes that rely on machine learning algorithms and highlight the challenges in designing efficient I/O transfers for such workflows. We comment on how to leverage the data access patterns in order to fetch in a more efficient way the required input data in the format and order given by the needs of the application and how to optimize the data path between collaborative processes. We will motivate our work and show performance gains with a study case of medical applications.

Gainaru, Ana↗

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA↗

Scalable Bayesian Physics-Informed Kolmogorov-Arnold Networks

Uncertainty quantification (UQ) plays a pivotal role in scientific machine learning, especially when surrogate models are used to approximate complex systems. Although multilayer perceptions (MLPs) are commonly employed as surrogates, they often suffer from overfitting due to their large number of parameters. Kolmogorov-Arnold networks (KANs) offer an alternative solution with fewer parameters. However, gradient-based inference methods, such as Hamiltonian Monte Carlo (HMC), may result in computational inefficiency when applied to KANs, especially for large-scale datasets, due to the high cost of back-propagation. To address these challenges, we propose a novel approach, combining the dropout Tikhonov ensemble Kalman inversion (DTEKI) with Chebyshev KANs. This gradient-free method effectively mitigates overfitting and enhances numerical stability. In addition, we incorporate the active subspace method to reduce the parameter-space dimensionality, allowing us to improve the accuracy of predictions and obtain more reliable uncertainty estimates. Extensive experiments demonstrate the efficacy of our approach in various test cases, including scenarios with large datasets and high noise levels. Our results show that the new method achieves comparable or better accuracy, much higher efficiency as well as stability compared to HMC, in addition to scalability. Moreover, by leveraging the low-dimensional parameter subspace, our method preserves prediction accuracy while substantially reducing further the computational cost.

97 MATHEMATICS AND COMPUTING↗

Automatic Generation of Algorithms for the Statistical Analysis of Planetary Nebulae Images

Analyzing data sets collected in experiments or by observations is a Core scientific activity. Typically, experimentd and observational data are &aught with uncertainty, and the analysis is based on a statistical model of the conjectured underlying processes, The large data volumes collected by modern instruments make computer support indispensible for this. Consequently, scientists spend significant amounts of their time with the development and refinement of the data analysis programs. AutoBayes [GF+02, FS03] is a fully automatic synthesis system for generating statistical data analysis programs. Externally, it looks like a compiler: it takes an abstract problem specification and translates it into executable code. Its input is a concise description of a data analysis problem in the form of a statistical model as shown in Figure 1; its output is optimized and fully documented C/C++ code which can be linked dynamically into the Matlab and Octave environments. Internally, however, it is quite different: AutoBayes derives a customized algorithm implementing the given model using a schema-based process, and then further refines and optimizes the algorithm into code. A schema is a parameterized code template with associated semantic constraints which define and restrict the template s applicability. The schema parameters are instantiated in a problem-specific way during synthesis as AutoBayes checks the constraints against the original model or, recursively, against emerging sub-problems. AutoBayes schema library contains problem decomposition operators (which are justified by theorems in a formal logic in the domain of Bayesian networks) as well as machine learning algorithms (e.g., EM, k-Means) and nu- meric optimization methods (e.g., Nelder-Mead simplex, conjugate gradient). AutoBayes augments this schema-based approach by symbolic computation to derive closed-form solutions whenever possible. This is a major advantage over other statistical data analysis systems which use numerical approximations even in cases where closed-form solutions exist. AutoBayes is implemented in Prolog and comprises approximately 75.000 lines of code. In this paper, we take one typical scientific data analysis problem-analyzing planetary nebulae images taken by the Hubble Space Telescope-and show how AutoBayes can be used to automate the implementation of the necessary anal- ysis programs. We initially follow the analysis described by Knuth and Hajian [KHO2] and use AutoBayes to derive code for the published models. We show the details of the code derivation process, including the symbolic computations and automatic integration of library procedures, and compare the results of the automatically generated and manually implemented code. We then go beyond the original analysis and use AutoBayes to derive code for a simple image segmentation procedure based on a mixture model which can be used to automate a manual preproceesing step. Finally, we combine the original approach with the simple segmentation which yields a more detailed analysis. This also demonstrates that AutoBayes makes it easy to combine different aspects of data analysis.

Fischer, Bernd↗

NeuroSEM: A hybrid framework for simulating multiphysics problems by coupling PINNs and spectral elements

Multiphysics problems that are characterized by complex interactions among fluid dynamics, heat transfer, structural mechanics, and electromagnetics, are inherently challenging due to their coupled nature. While experimental data on certain state variables may be available, integrating these data with numerical solvers remains a significant challenge. Physics-informed neural networks (PINNs) have shown promising results in various engineering disciplines, particularly in handling noisy data and solving inverse problems in partial differential equations (PDEs). However, their effectiveness in forecasting nonlinear phenomena in multiphysics regimes, particularly involving turbulence, is yet to be fully established. Here, this study introduces NeuroSEM, a hybrid framework integrating PINNs with the highfidelity Spectral Element Method (SEM) solver, Nektar++. NeuroSEM leverages the strengths of both PINNs and SEM, providing robust solutions for multiphysics problems. PINNs are trained to assimilate data and model physical phenomena in specific subdomains, which are then integrated into the Nektar++ solver. We demonstrate the efficiency and accuracy of NeuroSEM for thermal convection in cavity flow and flow past a cylinder. The framework effectively handles data assimilation by addressing those subdomains and state variables where the data is available. We applied NeuroSEM to the Rayleigh-B´enard convection system, including cases with missing thermal boundary conditions and noisy datasets. Finally, we applied the proposed NeuroSEM framework to real particle image velocimetry (PIV) data to capture flow patterns characterized by horseshoe vortical structures. Our results indicate that NeuroSEM accurately models the physical phenomena and assimilates the data within the specified subdomains. The framework’s plug-and-play nature facilitates its extension to other multiphysics or multiscale problems. Furthermore, NeuroSEM is optimized for efficient execution on emerging integrated GPU-CPU architectures. This hybrid approach enhances the accuracy and efficiency of simulations, making it a powerful tool for tackling complex engineering challenges in various scientific domains.

42 ENGINEERING↗

Understanding the Impact of Data Staging for Coupled Scientific Workflows

We report the rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows.

97 MATHEMATICS AND COMPUTING↗

Bringing chemical structures to life with augmented reality, machine learning, and quantum chemistry

Visualizing 3D molecular structures is crucial to understanding and predicting their chemical behavior. However, static 2D hand-drawn skeletal structures remain the preferred method of chemical communication. Here, we combine cutting-edge technologies in augmented reality (AR), machine learning, and computational chemistry to develop MolAR, an open-source mobile application for visualizing molecules in AR directly from their hand-drawn chemical structures. Users can also visualize any molecule or protein directly from its name or protein data bank ID and compute chemical properties in real time via quantum chemistry cloud computing. MolAR provides an easily accessible platform for the scientific community to visualize and interact with 3D molecular structures in an immersive and engaging way.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cooperative Research and Development Agreement With Georgetown University Report: National Institutes of Health, National Center for Advancing Translational Sciences Clinical and Translational Science Award

Oak Ridge National Laboratory (ORNL) is participating with Georgetown University (GU) as a subrecipient in response to the National Institutes of Health (NIH), National Center for Advancing Translational Sciences (NCATS) Clinical and Translational Science Award (CTSA) (U54 Clinical Trial Optional) funding opportunity announcement. This Cooperative Research and Development Agreement is put in to place to facilitate the development and implementation of clinical interventions that demonstrably improve human health is currently a complex, recursive, and inefficient process that leads to delays of years or decades before discoveries in biomedical research result in health benefits for patients and communities. NCATS conducts and supports research in the science of translation, to discover the mechanistic and operational principles of the intervention development and dissemination process, thereby providing the scientific foundation for improvements in translational efficiency that will accelerate the realization of interventions that improve human health. Under NCATS’ leadership, the CTSA Program supports a national network of medical research institutions called hubs. GU is the lead institution in one of the NIH hubs that was created as a result of a previous NIH CTSA. The missions of the GU have historically included the advancement of health through research in the clinical and biomedical sciences, the education of future leaders in medical and nursing practice and academia, and the provision of compassionate and scientifically competent patient care and service to the Washington, DC community and the nation. GU is the lead institution for the Georgetown-Howard Universities Center for Clinical and Translational Science (GHUCCTS), a multi-institutional partnership of medical research institutions forged from a desire to promote clinical research and translational science. Through multiple collaborations among these institutions, GHUCCTS is transforming clinical research and translational science in order to bring new scientific advances to health care. Oak Ridge National Laboratory is the Department of Energy's (DOE) largest science and energy laboratory. Managed since April 2000 by a partnership of the University of Tennessee and Battelle, ORNL was established in 1943 as a part of the secret Manhattan Project to pioneer a method for producing and separating plutonium. During the 1950s and 1960s, ORNL became an international center for the study of nuclear energy and related research in the physical and life sciences. With the creation of DOE in the 1970s, ORNL's mission broadened to include a variety of energy technologies and strategies. Today the laboratory supports the nation with a peacetime science and technology mission that is just as important as, but very different from, its role during the Manhattan Project. ORNL is home to the world's premier center for high performance supercomputing to enable scientific discovery. ORNL has extensive expertise in various areas of computer science that are uniquely situated to support GU. Additionally, ORNL’s leading computational user facilities present a unique opportunity to leverage the largest scale machines for open science in support of the stated mission of the NCATS CTSA. ORNL's partnership with GU will offer unparalleled opportunity in data analytics, deep-learning, artificial intelligence, and urban dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Convergent Manufacturing of Large-Scale Components for Nuclear Applications, via Additive Manufacturing and Powder Metallurgy Hot Isostatic Pressing

Powder metallurgy (PM)–hot isostatic pressing (PM-HIP) has long been recognized as a powerful route for producing fully dense, near net shape metallic components. By consolidating powders under high temperature and pressure, HIP provides isotropic properties, uniform microstructures, and scalability to complex geometries that are vital for sectors such as aerospace, energy, and nuclear power. Yet despite these advantages, the technology has remained constrained by costly trial and error canister fabrication, limitations of conventional forging, and incomplete knowledge about how the canister design influences final part properties. Additive manufacturing (AM), by contrast, thrives on design freedom and geometric flexibility but struggles with speed, scalability, and cost when applied to very large structures. The research presented in this report investigated how a convergent manufacturing approach, combining AM with PM-HIP, can merge the strengths of both technologies, leveraging AM’s flexibility for canister design and HIP’s consolidation capability to deliver reliable, large, and complex parts. The work progressed through three case studies that built on one another in scale and complexity. Small cylindrical canisters fabricated by conventional methods, laser powder bed fusion, and directed energy deposition were filled with stainless steel powders and subjected to HIP. The resulting parts demonstrated near-full density and mechanical properties on par with wrought stainless steel, showing for the first time that AM canisters can be a direct substitute for conventional ones without sacrificing quality. The next step involved a medium-scale, noncentrosymmetric T-valve, which is an enclosed, multibranch geometry that tested the limits of AM + PM-HIP integration. The T-valve achieved predictable shrinkage and uniform densification, confirming feasibility for enclosed designs. However, this study also revealed oxide inclusions and interfacial challenges at the AM + HIP boundary, underscoring the critical importance of controlling interface chemistry and employing robust, in situ strategies, such as melt pool monitoring and thermal monitoring, coupled with nondestructive evaluation techniques such as x-ray computed tomography. Finally, the effort culminated in fabricating a large-scale impeller weighing nearly 2000 lb and spanning 5 ft in diameter. Produced via multirobot wire arc AM and hot isostatic pressed to near-full density, the impeller validated industrial-scale feasibility. Predictive models closely matched experimental shrinkage, tensile properties were spatially uniform across the component, and the AM + PM-HIP interface proved mechanically sound despite the presence of oxide-decorated prior particle boundaries. This large-scale demonstration is a major milestone, showing that hybrid AM + PM‑HIP can reliably deliver components at reactor-relevant scales. Collectively, these studies charted a logical pathway: small-scale work built scientific confidence, medium-scale work highlighted opportunities and challenges, and large-scale work proved industrial impact. The overarching conclusion of this report is that AM + PM-HIP should not be seen as a replacement for forging but as a complementary pathway that provides the US with flexibility, resilience, and new options for manufacturing nuclear-grade components. Looking ahead, several directions emerge as critical to sustaining progress. Predictive modeling must become faster, more accessible, and more accurate, with digital twins and machine learning reducing reliance on trial and error. Powders and alloys must be optimized for HIP, with improved cleanliness, reduced oxides, and tailored chemistries that enhance creep, fatigue, and irradiation resistance. Interfaces between AM and HIP regions must be better engineered through coatings, machining strategies, and surface treatments to mitigate oxide formation and ensure reliable bonding to explore opportunities for HIP of targeted compositional parts, as well as multimaterial HIP cladding applications. Monitoring and nondestructive evaluation need to expand, incorporating multimodal sensors, x-ray computed tomography, and real-time data integration through platforms such as Pelican. At the same time, the pathway to industrial adoption requires techno-economic analysis, machinability studies, and qualification frameworks aligned with industry and regulatory standards. Finally, workforce and academic engagement must be strengthened. Programs that train technicians and engineers for US Navy and US Department of Energy manufacturing challenges should be paired with academic partnerships to support fundamental research, with open sharing of non-export-controlled data to accelerate innovation and build the next generation of experts. In conclusion, this report demonstrates that hybrid AM + PM-HIP is scientifically viable and strategically important. By combining the design agility of AM with the consolidation strength of HIP and embedding modeling, monitoring, and workforce development, this approach provided a transformative new capability for US manufacturing. The path forward is clear: hybrid AM + PM-HIP is not just a promising research direction but is also potentially an industrially relevant pathway that can reshape how nuclear-grade components are designed, qualified, and deployed.

36 MATERIALS SCIENCE↗

Machine Learning and Cosmology

Methods based on machine learning have recently made substantial inroads in many corners of cosmology. Through this process, new computational tools, new perspectives on data collection, model development, analysis, and discovery, as well as new communities and educational pathways have emerged. Despite rapid progress, substantial potential at the intersection of cosmology and machine learning remains untapped. In this white paper, we summarize current and ongoing developments relating to the application of machine learning within cosmology and provide a set of recommendations aimed at maximizing the scientific impact of these burgeoning tools over the coming decade through both technical development as well as the fostering of emerging communities.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Exploration of Quantum Machine Learning and AI Accelerators for Fusion Science

The dawn of the noisy intermediate-scale quantum (NISQ) era sparked rapid development in variational quantum algorithms. These algorithms, utilizing parameterized quantum circuit optimized by classical computers with feedback, are practical under the constraints of the current hardware and can potentially show quantum advantage. In recent years, these variational circuit are applied to neural networks, hoping to boost the triumphant success of deep learning. We simulate various of quantum-classical hybrid neural networks applied to scientific tasks, and seek to understand their capabilities and limitations. Specifically, we devise quantum convolutional layers and apply them to deep neural networks used for predicting plasma disruption in fusion reactors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Proxy App Suite Release (FY2020)

Version 4.0 of the ECP Proxy App Suite is practically unchanged from the previous release. The current set of proxies has proven useful for many aspects of benchmarking and co-design and we see little reason to alter the suite. Although there have been few changes to the ECP suite, the team has been hard at work in other areas. In the area of Machine Learning (ML) we have now created a separate proxy suite dedicated to this scientific applications of ML.

97 MATHEMATICS AND COMPUTING↗

Equation-based and data-driven modeling: Open-source software current state and future directions

Here, a review of current trends in scientific computing reveals a broad shift to open-source and higher-level programming languages such as Python and growing career opportunities over the next decade. Open-source modeling tools accelerate innovation in equation-based and data-driven applications. Significant resources have been deployed to develop data-driven tools (PyTorch, TensorFlow, Scikit-learn) from tech companies that rely on machine learning services to meet business needs while keeping the foundational tools open. Open-source equation-based tools such as Pyomo, CasADi, Gekko, and JuMP are also gaining momentum according to user community and development pace metrics. Integration of data-driven and principles-based tools is emerging. New compute hardware, productivity software, and training resources have the potential to radically accelerate progress. However, long-term support mechanisms are still necessary to sustain the momentum and maintenance of critical foundational packages.

97 MATHEMATICS AND COMPUTING↗

HPC-FAIR: A Framework Managing Data and AI Models for Analyzing and Optimizing Scientific Applications

The increasing reliance on machine learning (ML) to analyze and optimize large-scale scientific applications on supercomputers faces a significant bottleneck: the lack of readily available, high-quality training datasets and the difficulty in reusing existing AI models. This project was motivated by the urgent need to address the “FAIR” principles (Findability, Accessibility, Interoperability, Reusability) for both training datasets and AI models in the high-performance computing (HPC) domain. The project developed HPC-FAIR, a high-performance computing data management framework designed to centralize HPC-related datasets and AI models within a unified hub. To ensure interoperability, the framework established a standardized representation and vocabulary (ontology) for both data and models. HPC-FAIR also implemented automated workflows to streamline data processing, model access, and benchmarking. Additionally, the project focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics.

97 MATHEMATICS AND COMPUTING↗

ZENN: A thermodynamics-inspired computational framework for heterogeneous data–driven modeling

Traditional entropy-based methods—such as cross-entropy loss in classification problems—have long been essential tools for representing the information uncertainty and physical disorder in data and for developing artificial intelligence algorithms. However, the rapid growth of data across various domains has introduced new challenges, particularly the integration of heterogeneous datasets with intrinsic disparities. To address this, we introduce a zentropy-enhanced neural network (ZENN), extending zentropy theory into the data science domain via intrinsic entropy, enabling more effective learning from heterogeneous data sources. ZENN simultaneously learns both energy and intrinsic entropy components, capturing the underlying structure of multisource data. To support this, we redesign the neural network architecture to better reflect the intrinsic properties and variability inherent in diverse datasets. We demonstrate the effectiveness of ZENN on classification tasks and energy landscape reconstructions, showing its superior generalization capabilities and robustness-particularly in predicting high-order derivatives. In image and text classification tasks, ZENN demonstrates superior generalization by introducing a learnable temperature variable that models latent multisource heterogeneity, allowing it to surpass state-of-the-art models on CIFAR-10/100, BBC News, and AG News. As a practical application in materials science, we employ ZENN to reconstruct the Helmholtz energy landscape of Fe3Pt using data generated from density functional theory and capture key material behaviors, including negative thermal expansion and the critical point in the temperature–pressure space. Overall, this work presents a zentropy-grounded framework for data-driven machine learning, positioning ZENN as a versatile and robust approach for scientific problems involving complex, heterogeneous datasets.

36 MATERIALS SCIENCE↗