Text-to-Audio
Diffusers
English
text-video-to-audio
text-controlled-video-to-audio
audio-controlled-video-to-audio
audio-generation
Instructions to use YJX-Xiaomi/ControlFoley with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use YJX-Xiaomi/ControlFoley with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("YJX-Xiaomi/ControlFoley", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
YJX-Research commited on
Commit ·
0a021ca
1
Parent(s): 3d09e9b
Refresh full-task demo examples
Browse files- .gitattributes +2 -0
- README.md +9 -5
- examples/ac_v2a/acv2a_reference.wav +3 -0
- examples/ac_v2a/acv2a_video_input.mp4 +3 -0
- examples/ac_v2a/acv2a_video_output.mp4 +3 -0
- examples/ac_v2a/acv2a_video_output.wav +3 -0
- examples/t2a/prompt.txt +1 -0
- examples/t2a/t2a_output.wav +3 -0
- examples/tc_v2a/prompt.txt +1 -0
- examples/tc_v2a/tcv2a_video_input.mp4 +3 -0
- examples/tc_v2a/tcv2a_video_output.mp4 +3 -0
- examples/tc_v2a/tcv2a_video_output.wav +3 -0
- examples/tv2a/prompt.txt +1 -0
- examples/tv2a/tv2a_video_input.mp4 +3 -0
- examples/tv2a/tv2a_video_output.mp4 +3 -0
- examples/tv2a/tv2a_video_output.wav +3 -0
- examples/v2a/v2a_video_input.mp4 +3 -0
- examples/v2a/v2a_video_output.mp4 +3 -0
- examples/v2a/v2a_video_output.wav +3 -0
.gitattributes
CHANGED
|
@@ -39,3 +39,5 @@ assets/result1.png filter=lfs diff=lfs merge=lfs -text
|
|
| 39 |
assets/result2.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
assets/result3.png filter=lfs diff=lfs merge=lfs -text
|
| 41 |
assets/tease.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 39 |
assets/result2.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
assets/result3.png filter=lfs diff=lfs merge=lfs -text
|
| 41 |
assets/tease.png filter=lfs diff=lfs merge=lfs -text
|
| 42 |
+
examples/**/*.mp4 filter=lfs diff=lfs merge=lfs -text
|
| 43 |
+
examples/**/*.wav filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -299,34 +299,36 @@ Options:
|
|
| 299 |
|
| 300 |
### 📋 **Usage Examples**
|
| 301 |
|
|
|
|
|
|
|
| 302 |
- TV2A
|
| 303 |
|
| 304 |
```bash
|
| 305 |
-
python demo.py --video "
|
| 306 |
```
|
| 307 |
|
| 308 |
- TC-V2A
|
| 309 |
|
| 310 |
```bash
|
| 311 |
-
python demo.py --video "
|
| 312 |
```
|
| 313 |
|
| 314 |
- AC-V2A
|
| 315 |
|
| 316 |
```bash
|
| 317 |
-
python demo.py --video "
|
| 318 |
```
|
| 319 |
|
| 320 |
- V2A
|
| 321 |
|
| 322 |
```bash
|
| 323 |
-
python demo.py --video "
|
| 324 |
```
|
| 325 |
|
| 326 |
- T2A
|
| 327 |
|
| 328 |
```bash
|
| 329 |
-
python demo.py --prompt "A bird sings melodically in a forest." --duration
|
| 330 |
```
|
| 331 |
|
| 332 |
<hr style="border: none; border-top: 3px solid #333; margin: 16px 0;">
|
|
@@ -362,6 +364,8 @@ VGGSound, Kling-Audio-Eval, The Greatest Hits (<a href="https://creativecommons.
|
|
| 362 |
and MovieGen-Audio-Bench (<a href="https://creativecommons.org/licenses/by-nc/4.0/" target="_blank" style="color:#dc3545; text-decoration:none;">CC BY-NC 4.0</a>).<br>
|
| 363 |
All resources are used for <strong>academic and non-commercial demonstration purposes only</strong>.
|
| 364 |
|
|
|
|
|
|
|
| 365 |
This project is inspired by the following works:<br>
|
| 366 |
[stable-audio-tools](https://github.com/Stability-AI/stable-audio-tools), [MMAudio](https://github.com/hkchengrex/MMAudio), [Make-An-Audio 2](https://github.com/bytedance/Make-An-Audio-2), [Synchformer](https://github.com/v-iashin/Synchformer), and [audiocraft](https://github.com/facebookresearch/audiocraft).<br>
|
| 367 |
Thanks for their contributions.
|
|
|
|
| 299 |
|
| 300 |
### 📋 **Usage Examples**
|
| 301 |
|
| 302 |
+
Ready-to-use inputs and generated outputs for all supported tasks are available in [`examples/`](examples/). You can preview the results directly or reuse the inputs with the commands below.
|
| 303 |
+
|
| 304 |
- TV2A
|
| 305 |
|
| 306 |
```bash
|
| 307 |
+
python demo.py --video "examples/tv2a/tv2a_video_input.mp4" --prompt "skateboarding" --duration 8.0 --output "./output"
|
| 308 |
```
|
| 309 |
|
| 310 |
- TC-V2A
|
| 311 |
|
| 312 |
```bash
|
| 313 |
+
python demo.py --video "examples/tc_v2a/tcv2a_video_input.mp4" --prompt "thunder strike" --duration 6.0 --output "./output"
|
| 314 |
```
|
| 315 |
|
| 316 |
- AC-V2A
|
| 317 |
|
| 318 |
```bash
|
| 319 |
+
python demo.py --video "examples/ac_v2a/acv2a_video_input.mp4" --audio "examples/ac_v2a/acv2a_reference.wav" --duration 5.0 --output "./output"
|
| 320 |
```
|
| 321 |
|
| 322 |
- V2A
|
| 323 |
|
| 324 |
```bash
|
| 325 |
+
python demo.py --video "examples/v2a/v2a_video_input.mp4" --duration 6.0 --output "./output"
|
| 326 |
```
|
| 327 |
|
| 328 |
- T2A
|
| 329 |
|
| 330 |
```bash
|
| 331 |
+
python demo.py --prompt "A bird sings melodically in a forest." --duration 10.0 --output "./output"
|
| 332 |
```
|
| 333 |
|
| 334 |
<hr style="border: none; border-top: 3px solid #333; margin: 16px 0;">
|
|
|
|
| 364 |
and MovieGen-Audio-Bench (<a href="https://creativecommons.org/licenses/by-nc/4.0/" target="_blank" style="color:#dc3545; text-decoration:none;">CC BY-NC 4.0</a>).<br>
|
| 365 |
All resources are used for <strong>academic and non-commercial demonstration purposes only</strong>.
|
| 366 |
|
| 367 |
+
Demo media credits: audio from Pixabay; video from Pexels and Jimeng AI-generated content.
|
| 368 |
+
|
| 369 |
This project is inspired by the following works:<br>
|
| 370 |
[stable-audio-tools](https://github.com/Stability-AI/stable-audio-tools), [MMAudio](https://github.com/hkchengrex/MMAudio), [Make-An-Audio 2](https://github.com/bytedance/Make-An-Audio-2), [Synchformer](https://github.com/v-iashin/Synchformer), and [audiocraft](https://github.com/facebookresearch/audiocraft).<br>
|
| 371 |
Thanks for their contributions.
|
examples/ac_v2a/acv2a_reference.wav
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4556a5b298ccc593ebe930098e19d432a8f7a9e7d6f8447794bbbbe6ff7040b8
|
| 3 |
+
size 96846
|
examples/ac_v2a/acv2a_video_input.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d876b5d55e1887953f51bcdb6adf5c8c9d47f9f3080379cf03d9780b454ea811
|
| 3 |
+
size 4337622
|
examples/ac_v2a/acv2a_video_output.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6e746da1eaf7d00b5bef273e3c00c6e39117e75b9ee8d38922e918383c99b1b1
|
| 3 |
+
size 6332376
|
examples/ac_v2a/acv2a_video_output.wav
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0b8e2085d5af00c159f0f1277fc5d2471d53396919bb03a04694d5bbe3f0623a
|
| 3 |
+
size 444494
|
examples/t2a/prompt.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
A bird sings melodically in a forest
|
examples/t2a/t2a_output.wav
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:af3d770975ae5e65972c8c65303c7802fcecebb432d1b4de4f72ebb8d3f7c588
|
| 3 |
+
size 882766
|
examples/tc_v2a/prompt.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
thunder strike
|
examples/tc_v2a/tcv2a_video_input.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:1f4bb523bda7963bb2fa9b075ab95357a4ef326db4dc1ff312666831b6406e65
|
| 3 |
+
size 5051468
|
examples/tc_v2a/tcv2a_video_output.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7268363154f0a3895804b4075202f1df26c020fc354722fd41dde67ac1d2b6cb
|
| 3 |
+
size 7151147
|
examples/tc_v2a/tcv2a_video_output.wav
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c9d03c6463d775cdb20475308bd3f407a9d66ceffcd4c3a2d4d92fe07c13d7cb
|
| 3 |
+
size 1065038
|
examples/tv2a/prompt.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
skateboarding
|
examples/tv2a/tv2a_video_input.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:104e913e43ffb6bc6b218e1b4a11a52ae6021ade8dffb18240b3c9cf3439f0d0
|
| 3 |
+
size 6638507
|
examples/tv2a/tv2a_video_output.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c1940d5d07fbae6bf22a9c38a82017d06eed5117c4c39fbcd2496697defc3319
|
| 3 |
+
size 10184085
|
examples/tv2a/tv2a_video_output.wav
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:3f5d8e21a2dc67c2bf57affb0734738bbd54e22fd3cca357857bceb1dc6c9220
|
| 3 |
+
size 1417294
|
examples/v2a/v2a_video_input.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:03ab9996df934ae3fd1ec0294d81b665cd48ba3a534f9b1dd0def463caaee692
|
| 3 |
+
size 19530557
|
examples/v2a/v2a_video_output.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9817a06693b8893500cb605654a52fd2c1c5dd47e5b9274f1711dd3034fc0922
|
| 3 |
+
size 7824979
|
examples/v2a/v2a_video_output.wav
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9a25a79a63d30e980489b60a15235293bc7675224d94636c9496016c70a4b872
|
| 3 |
+
size 1065038
|