15 papers · ranked by Valyu relevance
Bowen Cui, Tejas Ramesh, Oscar Hernandez, Keren Zhou
Large Language Models (LLMs) have emerged as powerful tools for software development tasks such as code completion, translation, and optimization. However, their ability to generate efficient and correct code, particularly in complex High-Performance Computing (HPC) contexts, has remained underexplored. To address this…
Mohammad Zaeed, Tanzima Z. Islam, Vladimir Indic
Large language models (LLMs) show promise for automated code optimization. However, without performance context, they struggle to produce correct and effective code transformations. Existing performance tools can identify bottlenecks but stop short of generating actionable code changes. Consequently, performance…
Han Hu, Xiaoheng Xie, Bo Sun, Jian Gu + 2 more
Mobile apps frequently suffer from performance issues such as frame drops, overheating, and excessive power consumption. While developers optimize algorithms and debug code, a critical bottleneck often goes unnoticed: native libraries compiled with low optimization levels (O0/O1 instead of O2/O3). Because these…
Peter Arzt, Sebastian Kreutzer, Tim Jammer, Christian Bischof
Teaching performance engineering in high-performance computing (HPC) requires example codes that demonstrate bottlenecks and enable hands-on optimization. However, existing HPC applications and proxy apps often lack the balance of simplicity, transparency, and optimization potential needed for effective teaching. To…
Mohammed Fadle Abdulla
Developing an application with high performance through the code optimization places a greater responsibility on the programmers. While most of the existing compilers attempt to automatically optimize the program code, manual techniques remain the predominant method for performing optimization. Deciding where to try to…
Ke Cheng, Zhi Wang, Wen Hu, Tiannuo Yang + 2 more
—A service-level objective (SLO) is a target performance metric of service that cloud vendors aim to ensure. Delivering optimized SLOs can enhance user satisfaction and improve the competitiveness of cloud vendors. As large language models (LLMs) are gaining increasing popularity across various fields, it is of great…
Vibha Rajput, Alok Katiyar
—The aim of parallel computing is to increase an application's performance by executing the application on multiple processors. OpenMP is an API that supports multiplatform shared memory programming model and sharedmemory programs are typically executed by multiple threads. The use of multi threading can enhance the…
Elena Panova, Valentin Volokitin, Anton Gorshkov, Iosif Meyerov
> Abstract. The Black-Scholes option pricing problem is one of the widely used financial benchmarks. We explore the possibility of developing a high-performance portable code using the SYCL (Data Parallel C++) programming language. We start from a C++ code parallelized with OpenMP and show optimization techniques that…
E. H. Morel, Camille Coti
Performance portability is a major concern on current architectures. One way to achieve it is by using autotuning. In this paper, we are presenting how we exten ded a just-in-time compilation infrastructure to introduce autotuning capabiliti es triggered at run-time. When a function is executed, the first iterations…
Kyriakos Georgiou, Craig Blackmore, Samuel Xavier‐de‐Souza, Kerstin Eder
'Kerstin Eder'] Abstract. This paper presents the interesting observation that by performing fewer of the optimizations available in a standard compiler optimization level such as -O2, while preserving their original ordering, significant savings can be achieved in both execution time and energy consumption. This…
Lianjie Luo, Yang Chen, Chengyong Wu, Shun Long + 1 more
Iterative compilation is a widely adopted technique to optimize programs for different constraints such as performance, code size and power consumption in rapidly evolving hardware and software environments. However, in case of statically compiled programs, it is often restricted to optimizations for a specific dataset…
Chris Fawcett, Lars Kotthoff, Holger H. Hoos
Modern software systems in many application areas offer to the user a multitude of parameters, switches and other customisation hooks. Humans tend to have difficulties determining the best configurations for particular applications. Modern optimising compilers are an example of such software systems; their many…
E. A. Kiselev, P. N. Telegin, А. В. Баранов
> Abstract— The increase in performance and power of computing systems requires the wider use of program optimizations. The goal of performing optimizations is not only to reduce program runtime, but also to reduce other computer resources including power consumption. The goal of the study was to evaluate the impact of…
Joao B. Fernandes, Felipe H. S. da Silva, Samuel Xavier‐de‐Souza, Ítalo A. S. Assis
'Ítalo A. S. Assis'] Programs with high levels of complexity often face challenges in adjusting execution parameters, particularly when the ideal value for these parameters may change based on the execution context. These dynamic parameters significantly impact the program's performance. For instance, ideal parallel…
Jiří Filipovič, Jana Hozzová, Amin Nezarat, Jaroslav Oľha + 1 more
'Filip Petrović'] Nowadays, GPU accelerators are commonly used to speed up general-purpose computing tasks on a variety of hardware. However, due to the diversity of GPU architectures and processed data, optimization of codes for a particular type of hardware and specific data characteristics can be extremely…