Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Feature Engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Semi-supervised Hybrid Machine Learning Framework for the Qualification of Resistance Spot Welds

• Industries requiring high structural integrity, including automotive, aerospace, and construction, place considerable significance on weld quality classification. • The inspection normally involves human expertise through predefined quality metrics that are subjective, error-prone, and time-intensive • The challenge to classification model development is the scarcity of labeled data and imbalanced distributions in the data that are labeled. • This work develops a new hybrid methodology that achieves clustering using KMeans++ together with supervised classification to overcome these challenges. • The ensemble-based classifiers were identified as optimal, with accuracy enhancements of up to 8% using the pseudo-labeled dataset. • The work provides practical insight into feature engineering and machine learning integration in industrial quality assurance applications.

Rogers, Jeremy K. [Savannah River National Laborat↗

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

Unsupervised Detection of SOC Spoofing in OCPP 2.0.1 EV Charging Communication Protocol Using One-Class SVM

The electric vehicles (EVs) market keeps growing globally; thus, it is critical to secure the EV charging communication protocols in order to guarantee reliable and fair charging operations among the customers. The Open Charge Point Protocol (OCPP) 2.0.1 supports the communication between the Electric Vehicle Supply Equipment (EVSE) and Charging Station Management Systems (CSMSs); therefore, it becomes vulnerable to several types of attacks, which aim to jeopardize smart charging, billing, and energy management. Specifically, OCPP 2.0.1 allows the self-reporting of the State of Charge (SOC) values, which makes it vulnerable to spoofing-based cyberattacks, which target manipulating the scheduling priorities, distorting the load forecasts, and extending the charging sessions in an unfair manner. In this paper, we try to address this type of attack by providing a comprehensive analysis of the SOC spoofing attacks and introducing a novel unsupervised detection framework based on the One-Class Support Vector Machine (OCSVM) algorithm. Specifically, two types of attack scenarios are analyzed (i.e., priority manipulation and session extension) by deriving engineered features that capture the nonlinear relationships under normal charging behavior. Detailed simulation-based results are derived by utilizing the DESL-EPFL Level 3 EV charging dataset. Our results demonstrate high F1-score and recall in identifying spoofed SOC values and that the proposed OCSVM model demonstrates superior performance compared to alternative clustering and deep-learning based detectors.

EV charging↗

Data Augmentation for Neutron Spectrum Unfolding with Neural Networks

Neural networks require a large quantity of training spectra and detector responses in order to learn to solve the inverse problem of neutron spectrum unfolding. In addition, due to the under-determined nature of unfolding, non-physical spectra which would not be encountered in usage should not be included in the training set. While physically realistic training spectra are commonly determined experimentally or generated through Monte Carlo simulation, this can become prohibitively expensive when considering the quantity of spectra needed to effectively train an unfolding network. In this paper, we present three algorithms for the generation of large quantities of realistic and physically motivated neutron energy spectra. Using an IAEA compendium of 251 spectra, we compare the unfolding performance of neural networks trained on spectra from these algorithms, when unfolding real-world spectra, to two baselines. We also investigate general methods for evaluating the performance of and optimizing feature engineering algorithms.

McGreivy, James (ORCID:0000000321723411)↗

A Data-Driven Framework for Direct Local Tensile Property Prediction of Laser Powder Bed Fusion Parts

This article proposes a generalizable, data-driven framework for qualifying laser powder bed fusion additively manufactured parts using part-specific in situ data, including powder bed imaging, machine health sensors, and laser scan paths. To achieve part qualification without relying solely on statistical processes or feedstock control, a sequence of machine learning models was trained on 6299 tensile specimens to locally predict the tensile properties of stainless-steel parts based on fused multi-modal in situ sensor data and a priori information. A cyberphysical infrastructure enabled the robust spatial tracking of individual specimens, and computer vision techniques registered the ground truth tensile measurements to the in situ data. The co-registered 230 GB dataset used in this work has been publicly released and is available as a set of HDF5 files. The extensive training data requirements and wide range of size scales were addressed by combining deep learning, machine learning, and feature engineering algorithms in a relay. The trained models demonstrated a 61% error reduction in ultimate tensile strength predictions relative to estimates made without any in situ information. Lessons learned and potential improvements to the sensors and mechanical testing procedure are discussed.

36 MATERIALS SCIENCE↗

Open Data and Deep Semantic Segmentation for Automated Extraction of Building Footprints

Advances in machine learning and computer vision, combined with increased access to unstructured data (e.g., images and text), have created an opportunity for automated extraction of building characteristics, cost-effectively, and at scale. These characteristics are relevant to a variety of urban and energy applications, yet are time consuming and costly to acquire with today’s manual methods. Several recent research studies have shown that in comparison to more traditional methods that are based on features engineering approach, an end-to-end learning approach based on deep learning algorithms significantly improved the accuracy of automatic building footprint extraction from remote sensing images. However, these studies used limited benchmark datasets that have been carefully curated and labeled. How the accuracy of these deep learning-based approach holds when using less curated training data has not received enough attention. The aim of this work is to leverage the openly available data to automatically generate a larger training dataset with more variability in term of regions and type of cities, which can be used to build more accurate deep learning models. In contrast to most benchmark datasets, the gathered data have not been manually curated. Thus, the training dataset is not perfectly clean in terms of remote sensing images exactly matching the ground truth building’s foot-print. A workflow that includes data pre-processing, deep learning semantic segmentation modeling, and results post-processing is introduced and applied to a dataset that include remote sensing images from 15 cities and five counties from various region of the USA, which include 8,607,677 buildings. The accuracy of the proposed approach was measured on an out of sample testing dataset corresponding to 364,000 buildings from three USA cities. The results favorably compared to those obtained from Microsoft’s recently released US building footprint dataset.

97 MATHEMATICS AND COMPUTING↗

Recent Progress on Surface Water Quality Models Utilizing Machine Learning Techniques

Surface waterbodies are heavily exposed to pollutants caused by natural disasters and human activities. Empowering sensor technologies in water quality monitoring, sufficient measurements have become available to develop machine learning (ML) models. Numerous ML models have quickly been adopted to predict water quality indicators in various surface waterbodies. This paper reviews 78 recent articles from 2022 to October 2024, categorizing water quality models utilizing ML into three groups: Point-to-Point (P2P), which estimates the current target value based on other measurements at the same time point; Sequence-to-Point (S2P), which utilizes previous time series data to predict the target value at one time point ahead; and Sequence-to-Sequence (S2S), which uses previous time series data to forecast sequential target values in the future. The ML models used in each group are classified and compared according to water quality indicators, data availability, and model performance. Widely used strategies for improving performance, including feature engineering, hyperparameter tuning, and transfer learning, are recognized and described to enhance model effectiveness. The interpretability limitations of ML applications are discussed. This review provides a perspective on emerging ML for surface water quality models.

machine learning (ML)↗

What’s the Difference? The Potential for Convolutional Neural Networks for Transient Detection without Template Subtraction

Abstract We present a study of the potential for convolutional neural networks (CNNs) to enable separation of astrophysical transients from image artifacts, a task known as “real–bogus” classification, without requiring a template-subtracted (or difference) image, which requires a computationally expensive process to generate, involving image matching on small spatial scales in large volumes of data. Using data from the Dark Energy Survey, we explore the use of CNNs to (1) automate the real–bogus classification and (2) reduce the computational costs of transient discovery. We compare the efficiency of two CNNs with similar architectures, one that uses “image triplets” (templates, search, and difference image) and one that takes as input the template and search only. We measure the decrease in efficiency associated with the loss of information in input, finding that the testing accuracy is reduced from ∼96% to ∼91.1%. We further investigate how the latter model learns the required information from the template and search by exploring the saliency maps. Our work (1) confirms that CNNs are excellent models for real–bogus classification that rely exclusively on the imaging data and require no feature engineering task and (2) demonstrates that high-accuracy (>90%) models can be built without the need to construct difference images, but some accuracy is lost. Because, once trained, neural networks can generate predictions at minimal computational costs, we argue that future implementations of this methodology could dramatically reduce the computational costs in the detection of transients in synoptic surveys like Rubin Observatory's Legacy Survey of Space and Time by bypassing the difference image analysis entirely.

79 ASTRONOMY AND ASTROPHYSICS↗

An advanced regulator for the helium pressurization systems of the Space Shuttle OMS and RCS

The Space Shuttle Orbit Maneuvering System and Reaction Control System are pressure-fed rocket propulsion systems utilizing earth storable hypergolic propellants and featuring engines of 6000 lbs and 900 lbs thrust, respectively. The helium pressurization system requirements for these propulsion systems are defined and the current baseline pressurization systems are described. An advanced helium pressure regulator capable of meeting both OMS and RCS helium pressurization system requirements is presented and its operating characteristics and predicted performance characteristics are discussed.

Wichmann, H.↗

Results of acoustic testing of the JT8D-109 refan engines

A JT8D engine was modified to reduce jet noise levels by 6-8 PNdB at takeoff power without increasing fan generated noise levels. Designated the JT8D-109, the modified engines featured a larger single stage fan, and acoustic treatment in the fan discharge ducts. Noise levels were measured on an outdoor test facility for eight engine/acoustic treatment configurations. Compared to the baseline JT8D, the fully treated JT8D-109 showed reductions of 6 PNdB at takeoff, and 11 PNdB at a typical approach power setting.

Burdsall, E. A.↗

A flight instrumentation system for acquisition of atmospheric turbulence data

A flight instrumentation system for the acquisition of atmospheric turbulence data is described. Airflow direction transducers and an impact pressure transducer are the primary instruments for measuring vertical and lateral gust velocity, and a sensitive incremental pressure transducer is used to measure longitudinal gust velocity. Airplane motions, sensed by an inertial platform, are subtracted from the primary measurements during postflight data reduction to yield true gust velocity time histories. Salient engineering features of the instrumentation are discussed, and a complete description of the instrumentation is presented.

Meissner, C. W., Jr.↗

Space solar power systems

Studies were done on the feasibility of placing a solar power station called POwersat, in space. A general description of the engineering features are given as well as a brief discussion of the economic considerations.

Toliver, C.↗

Shuttle remote manipulator system workstation - Man-machine engineering

A major subsystem aboard the Shuttle Orbiter, the Remote Manipulator System (RMS) provides the capability to deploy and retrieve free-flying satellites, support attached payload operations, and aid in crewmember rescue from a disabled vehicle should the requirement arise. The Remote Manipulator System consists of 15.3-m (50 ft) articulated booms, end effectors, operator workstation, and closed circuit video, power and control subsystems. The manipulator booms (or arms) and end effectors are located in the Orbiter payload bay and operated from inside the cabin by one crewmember. This paper is primarily concerned with the design and development of the RMS operator's workstation, the man-machine engineering features and interfaces, and man-in-the-loop simulations and testing results obtained to date by the National Aeronautics and Space Administration (NASA).

Brown, J. W.↗

Engineering aspects of geothermal development with emphasis on the Imperial Valley of California

This review was prepared in support of a geothermal planning activity of the County of Imperial. Engineering features of potential geothermal development are outlined. Acreage requirements for drilling and powerplants are estimated, as are the costs for wells, fluid transmission pipes, and generating stations. Rough scaling relationships are developed for cost factors as a function of reservoir temperature. Estimates are made for cooling water requirements, and possible sources of cooling water are discussed. Availability and suitability of agricultural wastewater for cooling are emphasized. The utility of geothermal resources for fresh water production in the Imperial Valley is considered.

Goldsmith, M.↗

Life prediction modeling based on cyclic damage accumulation

A high temperature, low cycle fatigue life prediction method was developed. This method, Cyclic Damage Accumulation (CDA), was developed for use in predicting the crack initiation lifetime of gas turbine engine materials, where initiation was defined as a 0.030 inch surface length crack. A principal engineering feature of the CDA method is the minimum data base required for implementation. Model constants can be evaluated through a few simple specimen tests such as monotonic loading and rapic cycle fatigue. The method was expanded to account for the effects on creep-fatigue life of complex loadings such as thermomechanical fatigue, hold periods, waveshapes, mean stresses, multiaxiality, cumulative damage, coatings, and environmental attack. A significant data base was generated on the behavior of the cast nickel-base superalloy B1900+Hf, including hundreds of specimen tests under such loading conditions. This information is being used to refine and extend the CDA life prediction model, which is now nearing completion. The model is also being verified using additional specimen tests on wrought INCO 718, and the final version of the model is expected to be adaptable to most any high-temperature alloy. The model is currently available in the form of equations and related constants. A proposed contract addition will make the model available in the near future in the form of a computer code to potential users.

Nelson, Richard S.↗

Commercialization of dish-Stirling solar terrestrial systems

The requirements for dish-Stirling commercialization are described. The requirements for practical terrestrial power systems, both technical and economic, are described. Solar energy availability, with seasonal and regional variations, is discussed. The advantages and disadvantages of hybrid operation are listed. The two systems described use either a 25-kW free-piston Stirling hydraulic engine or a 5-kW kinematic Stirling engine. Both engines feature long-life characteristics that result from the use of welded metal bellows as hermetic seals between the working gas and the crankcase fluid. The advantages of the systems, the state of the technology, and the challenges that remain are discussed. Technology transfer between solar terrestrial Stirling applications and other Stirling applications is predicted to be important and synergistic.

Ross, Brad↗

Analyzing Thermal Conditions In Rocket Engines

Computer code, RTE, developed to perform three-dimensional thermal analyses of rocket thrust chambers. Calculates rate of heat transfer from combustion gases to coolant, coolant-temperature rise and pressure drop, and temperature profiles within cooling-jacket wall. Also calculates combustion-gas wall static pressure, temperature and enthalpy, as well as coolant pressure, temperature, and mach number for all stations. Program used for any propellant combination and most coolants commonly used in rockets. Code used for both regeneratively and radiatively cooled engines. However, in case of regeneratively cooled engines, applicability limited to engines featuring single-pass cooling and rectangular cooling channels.

Naraghi, M. H. N.↗

Quantitative Mapping of Reflected and Emitted Energy Patterns Over a City

There are major variations in energy flux within and across the region of a large city. These variations have impacts in disparate areas, such as human health, environmental monitoring and mitigation, and energy consumption. Knowledge of the variations also has utility to urban and regional planners, and climate modelers. The authors have developed a system which permits robust measurement of both the magnitude of the energy flux variation and the absolute value of energy flux over regions of the size of large cites. The technique uses properly acquired and processed multispectral imagery with bands in the visible, near-IR and thermal portions of the electromagnetic spectrum. With proper knowledge of the atmosphere and geometries of acquisition it is possible to compute the energy budget for individual pixels. The reality of this technique is demonstrated using data acquired over Salt Lake City, Utah. The deficiencies in the results emphasize the critical nature of various design and engineering features usually ignored in airborne and satellite imaging systems.

Luvall, J.↗