ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) | 979-8-3503-6874-1/25/$31.00 ©2025 IEEE | DOI: 10.1109/ICASSP49660.2025.10888189
Rhythmic Foley: A Framework For Seamless
Audio-Visual Alignment In Video-to-Audio
Synthesis
1 Zhiqi Huang∗‡ , 1 Dan Luo∗‡ , 2 Jun Wang† , 1 Huan Liao‡ , 1 Zhiheng Li† , 1 Zhiyong Wu†
1
Tsinghua University, Shenzhen, China
2
Tencent AI Lab, Shenzhen, China
{huangzq22, luod23}@mails.tsinghua.edu.cn
Abstract—Our research introduces an innovative framework
for video-to-audio synthesis, which solves the problems of audiovideo desynchronization and semantic loss in the audio. By
incorporating a semantic alignment adapter and a temporal
synchronization adapter, our method significantly improves semantic integrity and the precision of beat point synchronization,
particularly in fast-paced action sequences. Utilizing a contrastive
audio-visual pre-trained encoder, our model is trained with
video and high-quality audio data, improving the quality of the
generated audio. This dual-adapter approach empowers users
with enhanced control over audio semantics and beat effects,
allowing the adjustment of the controller to achieve better
results. Extensive experiments substantiate the effectiveness of
our framework in achieving seamless audio-visual alignment.
Index Terms—Video-to-audio Generation, Diffusion, Foley,
Audio-visual Alignment
I. I NTRODUCTION
In the current era of digital content creation, from filmmakers to social video creators, more and more people wish
to turn their creative visions into actual outputs. As text-tovideo models continue to improve and evolve, the demand for
adding vivid and fitting background audio to silent videos is
also increasing. To realize their creative ideas, creators need
a great deal of control. For social video creators, this control
should be simple yet powerful, while for filmmakers, it should
be precise and limitless. Our goal is to enable creators to
quickly iterate on their ideas and generate satisfactory audio
through simple controls.
Currently, there are two main methods for video-to-audio
generation, each with its advantages and disadvantages. The
first method [1], [2] builds on text-to-audio generation by
adding video event timestamps or video frame features to
control the alignment of semantics and sound events. The
second method [3], [4] directly trains a video-to-audio diffusion model. Although these methods can achieve video-toaudio generation, they still have three main limitations. (i)
Asynchronous. The first method is essentially still text-toaudio generation. Even with the addition of timestamps for
∗ Equal Contribution
alignment control, the instantaneous nature of sound events
means that current methods cannot ensure perfect alignment
between the generated audio and the actions in the video.
(ii) Low-quality. The second method uses the entire video as
input, aligning video and audio on the timeline, making more
efficient use of the video’s temporal and visual information.
However, due to the difficulty of obtaining high-quality videoaudio datasets, models trained with this method cannot guarantee the generation of high-quality audio. (iii) Semantic loss.
Additionally, some details may be lost during video encoding.
To overcome these limitations, we propose a new video-toaudio generation framework that adds two adapters to the pretrained video-to-audio model: a semantic alignment adapter
and a temporal synchronization adapter. Regarding semantic
loss, the semantic alignment adapter supplements the semantic
information and provides a simple audio event timestamp
to control the approximate time of semantic occurrences.
The semantic information generated by the large-scale visual
understanding model [5]–[7], can produce more comprehensive, temporally ordered descriptions compared to the simple
captions provided by the dataset. These descriptions include
more details and contextual information, helping the model to
better understand the content of the video. The timestamp can
help the model better align and associate information between
video and audio, thereby improving learning effectiveness.
To solve the problem of asynchronous, the temporal synchronization adapter provides the exact timing of instantaneous
audio events, significantly improving the temporal synchronization between audio and video.For the issue of poor audio
quality generated, we use the encoder aligned with audio
and video to perform visual sampling on pure audio data. We
use high-quality audio data for training the temporal adapter,
which greatly enhances the quality of audio generation.
The main contributions are summarized below:1
1) The semantic alignment adapter and temporal synchronization adapter can be used together or separately, enabling the generation of audio with complete semantic
content and finer temporal alignment.
‡ Work done during internship at Tencent AI Lab.
† Corresponding Author
1 Our demos are available at https://angelalilyer.github.io/RhythmicFoley/
Authorized licensed use limited to: Southern University of Science and Technology. Downloaded on August 18,2025 at 11:47:12 UTC from IEEE Xplore. Restrictions apply.
2) During our training process, we incorporated a substantial amount of high-quality pure audio data, which
significantly improved the quality and robustness of the
generated audio.
3) We propose a framework: Rhythmic Foley, which generates high-quality, fine-grained aligned audio for silent
videos. We conduct extensive quantitative and qualitative
experiments to verify the effectiveness of our method.
specifics of our model’s architecture, along with the procedures
for training and inference.
II. RELATED WORK
Text-to-Audio Generation The traditional Text-to-Audio
systems primarily rely on the signal processing methods and
machine learning to generate audio, which have limitations
in generation quality and computational costs. AudioLDM
[8], [9] improves these aspects by using pre-trained CLAP
model [10] to train latent diffusion models with audio embeddings, conditioned on text embeddings during sampling.
Some researchers [11], [12] have proposed data augmentation
strategies, which have improved the quality and robustness of
the audio generation. Other new researches [13], [14] have
also achieved better generation quality and generalization.
For example, Tango [15], [16], outperforms AudioLDM on
most metrics. This improvement is attributed to the adoption
of audio pressure level-based sound mixing for training set
augmentation. Baton [17] and Tango2 [16] curated a preference dataset with human and CLAP annotations respectively,
fine-tuning text-to-audio models via reinforcement learning
strategies. This achieved better text-audio alignment based on
both automatic and manual metrics.
Video-to-Audio Generation In recent years, the field of
video-to-audio generation has been extensively explored by researchers, who have developed various innovative approaches.
Some methods involve fine-tuning text-to-audio models and
integrating additional video information such as visual encodings or timestamps of video events to condition the
generation of audio. Examples of such approaches include
FoleyCrafter [1], SonicVisionLM [18], and V2A-Mapper [2].
Additionally, there are models [4], [19] that have been trained
from the ground up specifically for video-to-audio generation
using diffusion models. Simian et al. present Diff-Foley [3],
a synchronized Video-to-Audio synthesis method with a latent diffusion model that generates high-quality audio with
improved synchronization and audio-visual relevance. Beyond
these diffusion models, there are also video-to-audio generation methods [20], [21] that are based on large language
models. These leverage the semantic consistency between
video and audio, allowing for a mutual transmodal conversion.
III. METHOD
Providing a silent video, our approach can generate audio
that is both semantically aligned and precisely timed to match
key moments. Our framework is based on the video-to-audio
synthesis, where the audio-visual alignment encoder uses the
methodologies and weights from Diff-Foley [3]. For video
comprehension, we employ Mini-Gemini [5] to extract visual
content from the video. This section will delve into the
Fig. 1. An illustration of our framework during the training phase.
Fig. 2. An illustration of our framework during the inference phase.
A. Training-time Framework
As shown in Fig. 1, we input video or audio data into their
respective encoders to acquire aligned features that contain
visual and auditory information. The ground truth is also
encoded into a mel spectrum for subsequent processing. In
addition, we have designed two adapters for different training
conditions, which share the same architecture with the UNet
encoder [22] of the video-to-audio generator, which is the
same as ControlNet [23]. These adapters take the latent
embeddings from the UNet and extra conditional embeddings
as input and add their outputs as residuals to the UNet’s output.
During the training stage, only the adapter layers are trainable while the UNet blocks are frozen. The loss function used
is defined as:
L = Ez0 ,t,ct ,cf ,ϵ∼N (0,1) ||ϵ − ϵθ (zt , t, ct , cf )||22 ,
(1)
Authorized licensed use limited to: Southern University of Science and Technology. Downloaded on August 18,2025 at 11:47:12 UTC from IEEE Xplore. Restrictions apply.
where z0 represents the data in the latent space, ct and cf are
the video conditions such as semantics and timestamps and
the latent audio condition, respectively.
The Semantic Alignment Adapter uses text descriptions
extracted by Mini-Gemini [5] along with the outputs of event
detectors to enhance the semantic completeness and synchronicity of the generated audio. We qualitatively analyzed
some visual understanding models, by giving them silent
videos and assessing their ability to infer auditory events
in the videos. Mini-Gemini showed significant superiority in
the accuracy and completeness of the models’ descriptions
of auditory events. The video event detection is a simple
network [18], [24]trained to identify the occurrence of sound
events during a certain period. It helps models pay attention
to these visual cues, which is crucial for the transformation of
video content into audio. With the significant advancements
in large language models [25], [26], there has been substantial
progress in multimodal understanding [5], [27], [28]. They can
list possible sound events based on visual content, saving the
need for manual labeling. During the training of this module,
we have employed the VGGSound [29] and AudioSet [30]
datasets as our training data.
The Temporal Synchronization Adapter is designed to
synchronize the cadence of audio generation with the motion
in the video, employing a timestamp conditional mask to guide
the precise timing of sound effect generation. Derived from an
energy detector [31], this mask adeptly pinpoints onset events
in the audio stream. For the training of this module, we have
collected sound effect data to more accurately learn the rhythm
of the audio. The training data for this section was collected
from various high-quality audio websites.
B. Inference-time Framework
In the inference stage, as shown in 2, we have the option to
employ the semantic alignment adapter and temporal synchronization adapter either separately or in tandem with combined
weighting. For the temporal adapter, if users have higher
demands for precise dubbing, we can provide a visualization
annotation tool that can achieve better generation results. The
semantic adapter, on the other hand, obtains its conditional
input through the dialogue analysis of Mini-Gemini [5] and
detection by visual event detectors. Users can fine-tune the
strength of both adapters based on the actual generation
results, employing the framework to produce audio that is both
rhythmically precise and semantically comprehensive.
IV. E XPERIMENTS
In this section, we first introduce the experimental setup.
Then, we conducted experiments to compare our method
with previous approaches across several dimensions: semantic
alignment, temporal synchronization, and audio quality. Following this, we evaluate some objective and subjective metrics
to demonstrate the effectiveness and superiority of our model.
The specifics of these experiments and their corresponding
analyses will be outlined below.
TABLE I
Q UANTITATIVE RESULTS IN TERMS OF SEMANTIC ALIGNMENT
Testset
VGGSound
AVSync
Method
Diff-Foley [3]
V2A-Mapper [2]
FoleyCrafter [1]
Rhythmic-Foley(ours)
Diff-Foley [3]
Seeing and Hearing [32]
FoleyCrafter [1]
Rhythmic-Foley(ours)
FID↓
29.03
24.16
19.67
15.26
65.77
65.82
36.80
32.92
CLIP↑
9.172
9.720
10.70
20.82
10.38
2.033
11.94
18.92
MKL↓
3.318
2.654
2.561
1.684
1.963
2.547
1.497
1.524
A. Experimental Setup
During the training phase, our model was trained for 30
epochs on the VGGSound [29], AudioSet [30], and pure audio
dataset. During the evaluation phase, We conducted a comprehensive evaluation of our method, comparing it with FoleyCrafter [1], and techniques mentioned in its paper,including
Diff-Foley [3], V2A-Mapper [2], and Seeing-and-hearing [32].
We adhered strictly to the experimental setup outlined in the
article for our experiments. These evaluations were performed
on the AVSync15 [33] and VGGSound [29] datasets, employing a variety of metrics to assess semantic alignment
and audio quality, including Frechet Distance (FID) [34],
CLIP similarity [35], and Mean KL Divergence (MKL) [36].
Furthermore, we utilized onset detection accuracy (Onset Acc)
and onset detection average precision (Onset AP) [19] to
evaluate the synchronization accuracy of the generated audio.
B. Quantitative and Qualitative Comparison
Quantitative Comparison TABLE I presents a comparative
evaluation of our approach against state-of-the-art models in
terms of semantic alignment. We assessed three key metrics
multiple times on a subset of the VGGSound and AVSync15
test sets, and the table displays the mean values of these
tests. In our experiments, Rhythmic-Foley demonstrated superior performance on Frechet Distance and CLIP similarity.
Although the improvement in the MKL metric is not as
pronounced as in FID and CLIP similarity, the comparable
performance indicates that our method maintains a high level
of efficacy across different evaluation criteria. Additionally,
TABLE II illustrates the temporal synchronization outcomes
on the AVSync15 dataset. From these evaluations, it is evident
that our method excels in both semantic alignment with
visual prompts and temporal synchronization. Our approach
not only provides a more precise understanding of the audiovisual content but also generates audio that corresponds more
accurately with the video. This superior performance highlights the effectiveness of our model in achieving a more
coherent and synchronized audio-visual experience, setting a
new benchmark in the field.
Qualitative Comparison We provide the mel spectrogram
of the generated audio and ground in Fig. 3 for qualitative
comparison. As shown in Fig. 3, our model is able to extract
the correct semantics and the precise timing of events from the
video. Therefore, it achieves better semantic alignment and
Authorized licensed use limited to: Southern University of Science and Technology. Downloaded on August 18,2025 at 11:47:12 UTC from IEEE Xplore. Restrictions apply.
TABLE II
Q UANTITATIVE RESULTS IN TERMS OF TEMPORAL SYNCHRONIZATION
Method
Diff-Foley [3]
Seeing and Hearing [32]
FoleyCrafter [1]
Rhythmic-Foley(ours)
Onset AP↑
66.55
60.33
68.14
72.23
Onset ACC ↑
21.18
20.95
28.48
35.68
D. Ablation Study
To validate the efficacy of the semantic and temporal adapter
in the task of video-audio generation, we conducted a series
of ablation studies on the VGGSound and AVSync15 test sets.
Ablation on semantic adapter For this module, we conducted experiments using only event detection features, only
clip output features, and our method (event + clip) as inputs.
The TABLE IV shows that our method performs the best, as
the two features complement each other, enabling the model
to learn more accurate semantics from the representations.
TABLE IV
A BLATION ON SEMANTIC ADAPTER TESTED ON THE V GGSOUND
Method
w/o events
w/o semantic
Rhythmic-Foley(ours)
FID↓
15.26
18.38
15.06
CLIP↑
20.51
18.79
20.82
MKL↓
1.896
2.784
1.684
Ablation on temporal adapter As illustrated in TABLE
V, the first method employs video features aligned with
audio as input features for training, which yields results
comparable to our approach (where audio features aligned
with video are used for training), albeit slightly inferior, with
a negligible difference. However, models lacking a temporal
adapter possess only semantic control information without
frame-level timestamp details, leading to a marked decrease
in the accuracy of time synchronization and the audio quality.
Fig. 3. Mel spectrogram comparision
higher temporal synchronization compared to other models.
Without the temporal adapter, our model can detect events
occurring in the video and predict the coarse occurrence
and duration of these events through the semantic alignment
adapter. The complete Rhythmic Foley, with the help of the
temporal synchronization adapter, can also predict more accurate and precise timings of instantaneous actions. Therefore,
Rhythmic Foley can achieve SOTA in these aspects.
TABLE V
A BLATION ON TEMPORAL ADAPTER TESTED ON THE AVS YNC 15.
Method
trained by videos
w/o temporal
Rhythmic-Foley(ours)
Onset AP↑
70.31
66.33
72.23
Onset ACC ↑
34.18
21.95
35.68
VISQOL ↑
3.175
3.117
3.215
V. C ONCLUSION
TABLE III
AVERAGE U SER R ANKING (AUR) IN TERMS OF AUDIO QUALITY,
SEMANTIC ALIGNMENT, AND TEMPORAL SYNCHRONIZATION .
Method
Diff-Foley [3]
FoleyCrafter [1]
Rhythmic-Foley(ours)
Quality↑
3.283
3.387
3.951
Semantic↑
3.467
3.868
3.917
Temporal↑
3.260
2.587
4.236
C. User Study
We further invited users to evaluate the audio quality, semantic consistency, and temporal alignment of audio generated
by different methods. The Average Human Ranking (AHR)
was employed to rate each result on a scale of 1 to 5 (with
higher scores indicating better performance). Respondents
rated a set of audio clips generated by different methods for the
same video. We collected 2025 responses, which are presented
in TABLE III. Our method achieved the best in the subjective
evaluation across three key aspects.
In this paper, we introduce a novel framework for videoto-audio generation, which integrates two adapters into the
pre-trained video-to-audio generation model: the semantic
alignment adapter and the temporal synchronization adapter.
These adapters address the issues of semantic loss in audio
generation and the asynchrony between audio and video actions, respectively. Furthermore, our training process involves
mixing audio and video data, employing an audio encoder that
is aligned with the video encoder. Through extensive quantitative experiments and user studies, we have demonstrated the
effectiveness of our method, which significantly enhances the
quality of video-to-audio generation.
ACKNOWLEDGMENT
This work is supported by National Natural Science Foundation of China (62076144) and Shenzhen Science and Technology Program (WDZC20220816140515001,
JCYJ20220818101014030).
Authorized licensed use limited to: Southern University of Science and Technology. Downloaded on August 18,2025 at 11:47:12 UTC from IEEE Xplore. Restrictions apply.
R EFERENCES
[1] Yiming Zhang, Yicheng Gu, Yanhong Zeng, Zhening Xing, Yuancheng
Wang, Zhizheng Wu, and Kai Chen, “Foleycrafter: Bring silent
videos to life with lifelike and synchronized sounds,” arXiv preprint
arXiv:2407.01494, 2024.
[2] Heng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright, and
Weidong Cai, “V2a-mapper: A lightweight solution for vision-to-audio
generation by connecting foundation models,” in Proceedings of the
AAAI Conference on Artificial Intelligence, 2024, vol. 38, pp. 15492–
15501.
[3] Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao, “Diff-foley:
Synchronized video-to-audio synthesis with latent diffusion models,”
Advances in Neural Information Processing Systems, vol. 36, 2024.
[4] Manjie Xu, Chenxing Li, Yong Ren, Rilin Chen, Yu Gu, Wei Liang, and
Dong Yu, “Video-to-audio generation with hidden alignment,” arXiv
preprint arXiv:2407.07464, 2024.
[5] Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin
Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia, “Mini-gemini: Mining
the potential of multi-modality vision language models,” arXiv preprint
arXiv:2403.18814, 2024.
[6] Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang
Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, et al.,
“Sharegpt4video: Improving video understanding and generation with
better captions,” arXiv preprint arXiv:2406.04325, 2024.
[7] Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian,
Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang,
et al., “Internlm-xcomposer-2.5: A versatile large vision language
model supporting long-contextual input and output,” arXiv preprint
arXiv:2407.03320, 2024.
[8] Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo
Mandic, Wenwu Wang, and Mark D Plumbley, “Audioldm: Textto-audio generation with latent diffusion models,” arXiv preprint
arXiv:2301.12503, 2023.
[9] Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian,
Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D Plumbley,
“Audioldm 2: Learning holistic audio generation with self-supervised
pretraining,” IEEE/ACM Transactions on Audio, Speech, and Language
Processing, 2024.
[10] Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang, “Clap learning audio concepts from natural language
supervision,” in ICASSP 2023-2023 IEEE International Conference on
Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp.
1–5.
[11] Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor BergKirkpatrick, and Shlomo Dubnov, “Large-scale contrastive languageaudio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP 2023-2023 IEEE International Conference on
Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp.
1–5.
[12] Yi Yuan, Haohe Liu, Xubo Liu, Qiushi Huang, Mark D Plumbley,
and Wenwu Wang, “Retrieval-augmented text-to-audio generation,” in
ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech
and Signal Processing (ICASSP). IEEE, 2024, pp. 581–585.
[13] Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu,
Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William
Ngan, et al., “Audiobox: Unified audio generation with natural language
prompts,” arXiv preprint arXiv:2312.15821, 2023.
[14] Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov,
Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David
Grangier, Marco Tagliasacchi, et al., “Audiolm: a language modeling
approach to audio generation,” IEEE/ACM transactions on audio,
speech, and language processing, vol. 31, pp. 2523–2533, 2023.
[15] Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya
Poria, “Text-to-audio generation using instruction-tuned llm and latent
diffusion model,” arXiv preprint arXiv:2304.13731, 2023.
[16] Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu,
Rada Mihalcea, and Soujanya Poria, “Tango 2: Aligning diffusion-based
text-to-audio generations through direct preference optimization,” arXiv
preprint arXiv:2404.09956, 2024.
[17] Huan Liao, Haonan Han, Kai Yang, Tianjiao Du, Rui Yang, Zunnan Xu,
Qinmei Xu, Jingquan Liu, Jiasheng Lu, and Xiu Li, “Baton: Aligning
text-to-audio model with human preference feedback,” arXiv preprint
arXiv:2402.00744, 2024.
[18] Zhifeng Xie, Shengye Yu, Qile He, and Mengtian Li, “Sonicvisionlm:
Playing sound with vision language models,” in Proceedings of the
IEEE/CVF Conference on Computer Vision and Pattern Recognition,
2024, pp. 26866–26875.
[19] Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, and Andrew
Owens, “Conditional generation of audio from video via foley analogies,” in Proceedings of the IEEE/CVF Conference on Computer Vision
and Pattern Recognition, 2023, pp. 2426–2436.
[20] Gehui Chen, Guan’an Wang, Xiaowen Huang, and Jitao Sang, “Semantically consistent video-to-audio generation using multimodal language
large model,” arXiv preprint arXiv:2404.16305, 2024.
[21] Artemis Panagopoulou, Le Xue, Ning Yu, Junnan Li, Dongxu Li, Shafiq
Joty, Ran Xu, Silvio Savarese, Caiming Xiong, and Juan Carlos Niebles,
“X-instructblip: A framework for aligning x-modal instruction-aware
representations to llms and emergent cross-modal reasoning,” arXiv
preprint arXiv:2311.18799, 2023.
[22] Olaf Ronneberger, Philipp Fischer, and Thomas Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical
image computing and computer-assisted intervention–MICCAI 2015:
18th international conference, Munich, Germany, October 5-9, 2015,
proceedings, part III 18. Springer, 2015, pp. 234–241.
[23] Lvmin Zhang, Anyi Rao, and Maneesh Agrawala, “Adding conditional
control to text-to-image diffusion models,” in Proceedings of the
IEEE/CVF International Conference on Computer Vision, 2023, pp.
3836–3847.
[24] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Identity
mappings in deep residual networks,” in Computer Vision–ECCV 2016:
14th European Conference, Amsterdam, The Netherlands, October 11–
14, 2016, Proceedings, Part IV 14. Springer, 2016, pp. 630–645.
[25] Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge, Ying Shan, and Ge Li,
“St-llm: Large language models are effective temporal learners,” arXiv
preprint arXiv:2404.00308, 2024.
[26] Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li,
Zejun Ma, and Chunyuan Li, “Llava-next-interleave: Tackling multiimage, video, and 3d in large multimodal models,” arXiv preprint
arXiv:2407.07895, 2024.
[27] Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with
advanced large language models,” arXiv preprint arXiv:2304.10592,
2023.
[28] Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan
Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong,
and Mohamed Elhoseiny, “Minigpt-v2: large language model as a
unified interface for vision-language multi-task learning,” arXiv preprint
arXiv:2310.09478, 2023.
[29] Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman,
“Vggsound: A large-scale audio-visual dataset,” in ICASSP 20202020 IEEE International Conference on Acoustics, Speech and Signal
Processing (ICASSP). IEEE, 2020, pp. 721–725.
[30] Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen,
Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter,
“Audio set: An ontology and human-labeled dataset for audio events,”
in 2017 IEEE international conference on acoustics, speech and signal
processing (ICASSP). IEEE, 2017, pp. 776–780.
[31] Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt
McVicar, Eric Battenberg, and Oriol Nieto, “librosa: Audio and music
signal analysis in python.,” in SciPy, 2015, pp. 18–24.
[32] Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang, and Qifeng
Chen, “Seeing and hearing: Open-domain visual-audio generation with
diffusion latent aligners,” in Proceedings of the IEEE/CVF Conference
on Computer Vision and Pattern Recognition, 2024, pp. 7151–7161.
[33] Lin Zhang, Shentong Mo, Yijing Zhang, and Pedro Morgado, “Audiosynchronized visual animation,” arXiv preprint arXiv:2403.05659, 2024.
[34] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard
Nessler, and Sepp Hochreiter, “Gans trained by a two time-scale
update rule converge to a local nash equilibrium,” Advances in neural
information processing systems, vol. 30, 2017.
[35] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel
Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin,
Jack Clark, et al., “Learning transferable visual models from natural
language supervision,” in International conference on machine learning.
PMLR, 2021, pp. 8748–8763.
[36] Vladimir Iashin and Esa Rahtu, “Taming visually guided sound
generation,” arXiv preprint arXiv:2110.08791, 2021.
Authorized licensed use limited to: Southern University of Science and Technology. Downloaded on August 18,2025 at 11:47:12 UTC from IEEE Xplore. Restrictions apply.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )