16 papers · ranked by Valyu relevance
Paul Cardosi, Bérenger Bramas, Bilal Alatas
Parallelization is needed everywhere, from laptops and mobile phones to supercomputers. Among parallel programming models, task-based programming has demonstrated a powerful potential and is widely used in high-performance scientific computing. Not only does it allow efficient parallelization across distributed…
Oleksii Oleksenko, Bohdan Trach, Mark Silberstein, Christof Fetzer
SpecFuzz is the first tool that enables dynamic testing for speculative execution vulnerabilities (e.g., Spectre). The key is a novel concept of speculation exposure: The program is instrumented to simulate speculative execution in software by forcefully executing the code paths that could be triggered due to…
Bérenger Bramas, Gang Mei
Task-based programming models have demonstrated their efficiency in the development of scientific applications on modern high-performance platforms. They allow delegation of the management of parallelization to the runtime system (RS), which is in charge of the data coherency, the scheduling, and the assignment of the…
Giorgi Maisuradze, Christian Rossow
—Whenever modern CPUs encounter a conditional branch for which the condition cannot be evaluated yet, they predict the likely branch target and speculatively execute code. Such pipelining is key to optimizing runtime performance and is incorporated in CPUs for more than 15 years. In this paper, to the best of our…
Esmaeil Mohammadian Koruyeh, Khaled N. Khasawneh, Chengyu Song, Nael Abu‐Ghazaleh
'Nael Abu‐Ghazaleh'] The recent Spectre attacks exploit speculative execution, a pervasively used feature of modern microprocessors, to allow the exfiltration of sensitive data across protection boundaries. In this paper, we introduce a new Spectreclass attack that we call SpectreRSB. In particular, rather than…
Alvaro Estebanez, Diego R. Llanos, David Orden, Belen Palop + 1 more
'Rafael Sachetto Oliveira'] Loops are a rich source of parallelism. Unfortunately, many loops cannot be safely parallelized at compile time because the compiler is not able to guarantee that there will be no dependence violations. Thread-Level Speculation (TLS) techniques, either hardware or software-based, allow the…
Khaled N. Khasawneh, Esmaeil Mohammadian Koruyeh, Chengyu Song, Dmitry Evtyushkin + 2 more
'Dmitry Evtyushkin' 'Dmitry Ponomarev' 'Nael Abu‐Ghazaleh'] Abstract—Speculative execution, which is used pervasively in modern CPUs, can leave side effects in the processor caches and other structures even when the speculated instructions do not commit and their direct effect is not visible. The recent Meltdown and…
Taposh Dutta-Roy
—This paper is a review of the developments in Instruction level parallelism. It takes into account all the changes made in speeding up the execution. The various drawbacks and dependencies due to pipelining are discussed and various solutions to overcome them are also incorporated. It goes ahead in the last section to…
Ali Sahraee
—Spectre attacks exploit speculative execution to leak sensitive information. In the last few years, a number of static side-channel detectors have been proposed to detect cache leakage in the presence of speculative execution. However, these techniques either ignore branch prediction mechanism, detect static…
Fan Xu, Li Shen, Zhiying Wang, Bo Su + 2 more
Exploiting potential thread-level parallelism (TLP) is becoming the key factor to improving performance of programs on multicore or many-core systems. Among various kinds of parallel execution models, the software-based speculative parallel model has become a research focus due to its low cost, high efficiency…
Rasha Omar, Mostafa Abbas, Ahmed El-Mahdy, Erven Rohou + 1 more
'Rafael Sachetto Oliveira'] With the widespread of multicore systems, automatic parallelization becomes more pronounced, particularly for legacy programs, where the source code is not generally available. An essential operation in any parallelization system is detecting data dependence among parallelization candidate…
Kimonas Provatas, Aris Karatzikos, Charalampos Koilakos, Michail Patsakis + 5 more
Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose…
Hao Ding, Nannan Wu, Tianyi Qiu
DNA foundation models such as Evo2 7B adopt hybrid Hyena/attention architectures (Striped-Hyena2) whose single-stream autoregressive decoding is bounded by weight bandwidth at ∼45 tok/s. Speculative decoding on such hybrids faces a systems problem that prior SSM work solves only partially: after a draft is verified…
Natnatee Dokmai, Can Kockan, Kaiyuan Zhu, XiaoFeng Wang + 2 more
Genotype imputation is an essential tool in genetics research, whereby missing genotypes are inferred based on a panel of reference genomes to enhance the power of downstream analyses. Recently, public imputation servers have been developed to allow researchers to leverage increasingly large-scale and diverse genetic…
Olivia A. Kim, Alexander D. Forrence, Samuel D. McDougle
Prediction errors guide many forms of learning, providing teaching signals that help us improve our performance. Implicit motor adaptation, for instance, is driven by sensory prediction errors (SPEs), which occur when the expected and observed consequences of a movement differ. Traditionally, SPE computation is thought…
Anand, Vaastav, Garg, Deepak + 2 more
Specializing systems to specifics of the workload they serve and platform they are running on often significantly improves performance. However, specializing systems is difficult in practice because of compounding challenges: i) complexity for the developers to determine and implement optimal specialization; ii)…