Search · four archives
Search · four archives
21 papers · ranked by Valyu relevance
Mengshun Hu, Kui Jiang, Zhixiang Nie, Zheng Wang
Spatial-Temporal Video Super-Resolution (ST-VSR) technology generates high-quality videos with higher resolution and higher frame rates. Existing advanced methods accomplish ST-VSR tasks through the association of Spatial and Temporal video super-resolution (S-VSR and T-VSR). These methods require two alignments and…
Junpeng Jing, Mao Ye, Krystian Mikolajczyk
Stereo Matching Authors: ['Junpeng Jing' 'Mao Ye' 'Krystian Mikolajczyk'] > Abstract. Dynamic stereo matching is the task of estimating consistent disparities from stereo videos with dynamic objects. Recent learningbased methods prioritize optimal performance on a single stereo pair, resulting in temporal…
Jingbei Li, Meng Yi, Zhiyong Wu, Helen Meng + 3 more
'Yuping Wang' 'Yuxuan Wang'] Although deep learning and end-to-end models have been widely used and shown their superiority in automatic speech recognition (ASR) and text-to-speech (TTS) synthesis, state-of-the-art forced alignment (FA) models are still based on hidden Markov model (HMM). HMM has limited view of…
Jinxia Yang, Bing Su, Wayne Xin Zhao, Ji-Rong Wen
Multimodal Pre-training Authors: ['Jinxia Yang' 'Bing Su' 'Wayne Xin Zhao' 'Ji-Rong Wen'] Medical vision-language pre-training methods mainly leverage the correspondence between paired medical images and radiological reports. Although multi-view spatial images and temporal sequences of image-report pairs are available…
Mahmoud Rokaya, Dalia I. Hemdan, Mohammed A. Alzain, El-Sayed Atlam
Introduction A central limitation of existing temporal image analysis and video understanding models lies in their reliance on explicit motion cues, dense supervision, or auxiliary modalities, which constrains their ability to infer latent temporal structure, evolving semantic states, and long-range dependencies from…
Wisnu Aditya, Timothy K. Shih, Tipajin Thaipisutikul, Arda Satata Fitriajie + 5 more
'Arda Satata Fitriajie' 'Munkhjargal Gochoo' 'Fitri Utaminingrum' 'Chih-Yang Lin' 'Loris Nanni' 'Leon Rothkrantz'] Given video streams, we aim to correctly detect unsegmented signs related to continuous sign language recognition (CSLR). Despite the increase in proposed deep learning methods in this area, most of them…
Ran Armoni, Elhanan Borenstein
A major challenge in working with longitudinal data when studying some temporal process is the fact that differences in pace and dynamics might overshadow similarities between processes. In the case of longitudinal microbiome data, this may hinder efforts to characterize common temporal trends across individuals or to…
Santiago Marco-Sola, Jordan M. Eizenga, Andrea Guarracino, Benedict Paten + 2 more
Pairwise sequence alignment remains a fundamental problem in computational biology and bioinformatics. Recent advances in genomics and sequencing technologies demand faster and scalable algorithms that can cope with the ever-increasing sequence lengths. Classical pairwise alignment algorithms based on dynamic…
Ziyang Song, Qincheng Lu, Zhu He, Yue Li
Representation Learning Authors: ['Ziyang Song' 'Qincheng Lu' 'Zhu He' 'Yue Li'] Learning time-series representations for discriminative tasks has been a long-standing challenge. Current pre-training methods are limited in either unidirectional next-token prediction or randomly masked token prediction. We propose a…
Rongpei Gou, Jingyi Yang, Menghan Guo, Yingjun Chen + 1 more
Central nervous system (CNS) drugs have had a significant impact on human health, e.g., treating a wide range of neurodegenerative and psychiatric disorders. In recent years, deep learning-based generative models, particularly those for designing drugs from scratch, have shown great potential for accelerating drug…
Christina Sartzetaki, Anne W. Zonneveld, Pablo Oyarzo, Alessandro T. Gifford + 3 more
The human brain is the most efficient and versatile system for processing dynamic visual input. By comparing representations from deep video models to brain activity, we can gain insights into mechanistic solutions for effective video processing, important to better understand the brain and to build better models.…
Xiaowei Han, Wenbao Si, Honghui Zhang, Maolin Yang + 4 more
Fine-grained video action recognition remains challenging because action categories often differ only in subtle inter-class variations and complex temporal dynamics. Recent Contrastive Language-Image Pre-training (CLIP)-based extensions perform well on general action recognition, but they typically rely on early global…
Ming Xu, Sourav Garg, Michael Milford, Stephen Jay Gould
This paper addresses learning end-to-end models for time series data that include a temporal alignment step via dynamic time warping (DTW). Existing approaches to differentiable DTW either differentiate through a fixed warping path or apply a differentiable relaxation to the min operator found in the recursive steps…
Yoav Shalev, Lior Wolf
We study the problem of syncing the lip movement in a video with the audio stream. Our solution finds an optimal alignment using a dual-domain recurrent neural network that is trained on synthetic data we generate by dropping and duplicating video frames. Once the alignment is found, we modify the video in order to…
Kenan Li, Katherine Sward, Huiyu Deng, John Morrison + 6 more
Advances in measurement technology are producing increasingly time-resolved environmental exposure data. We aim to gain new insights into exposures and their potential health impacts by moving beyond simple summary statistics (e.g., means, maxima) to characterize more detailed features of high-frequency time series…
Yiqi Jiang, Kaiwen Sheng, Yujia Gao, E. Kelly Buchanan + 7 more
Recent work indicates that low-dimensional dynamics of neural and behavioral data are often preserved across days and subjects. However, extracting these preserved dynamics remains challenging: high-dimensional neural population activity and the recorded neuron populations vary across recording sessions. While existing…
Frantisek Forgac, Dasa Munkova, Michal Munk, Livia Kelebercova
Parallel texts represent a very valuable resource in many applications of natural language processing. The fundamental step in creating parallel corpus is the alignment. Sentence alignment is the issue of finding correspondence between source sentences and their equivalent translations in the target text. A number of…
Jiaming Xu, Tien Dung Nguyen, Jerry Tang, Alexander G. Huth + 1 more
Large language models trained on next-word prediction have impressive linguistic capabilities. This suggests that the goal of temporal prediction is essential to language processing, but how this goal impacts the structure of speech representations in the human brain remains unknown. Here, we test the hypothesis that…
Max Doblas, Oscar Lostes-Cazorla, Quim Aguado-Puig, Cristian Iñiguez + 2 more
Pairwise sequence alignment is a core component of multiple sequencing-data analysis tools. Recent advancements in sequencing technologies have enabled the generation of longer sequences at a much lower price. Thus, long-read sequencing technologies have become increasingly popular in sequencing-based studies. However…
Authors not listed
This research presents a novel approach to obstacle detection during navigation using a combination of Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks. The primary objective is to generate accurate image captions that describe the content of images, which is crucial for applications such…
Yangwen Xu, Nicola Sartorato, Léo Dutriaux, Roberto Bottini
Humans conceptualize time in terms of space, allowing flexible time construals from various perspectives. We can travel internally through a timeline to remember the past and imagine the future (i.e., mental time travel) or watch from an external standpoint to have a panoramic view of history (i.e., mental time…