5 2

Wok

woctezuma

AI & ML interests

🦆

Recent Activity

liked a Space 23 days ago

black-forest-labs/FLUX.1-Kontext-Dev

liked a Space about 2 months ago

linoyts/scribble-sdxl-flash

reacted to sayakpaul's post with 🔥 12 months ago

Flux.1-Dev like images but in fewer steps. Merging code (very simple), inference code, merged params: https://huggingface.co/sayakpaul/FLUX.1-merged Enjoy the Monday 🤗

View all activity

Organizations

liked a Space 23 days ago

938

FLUX.1 Kontext

⚡

Kontext image editing on FLUX[dev]

liked a Space about 2 months ago

179

Flash Scribble SDXL

⚡

Generate art from sketches or images

reacted to sayakpaul's post with 🔥 12 months ago

Post

4544

Flux.1-Dev like images but in fewer steps.

Merging code (very simple), inference code, merged params: sayakpaul/FLUX.1-merged

Enjoy the Monday 🤗

4 replies

reacted to gokaygokay's post with 👍🔥 12 months ago

Post

4707

I've made a creative version of Tile Upscaler

- gokaygokay/TileUpscalerV2

- https://github.com/gokayfem/Tile-Upscaler

- New tiling strategy
- Now it's closer to Clarity Upscaler
- It has more parameters to play and it has more room to fail because of that
- You should try different resolutions, strength and controlnet strength

Original Tile Upscaler
- gokaygokay/Tile-Upscaler

reacted to gokaygokay's post with 👍 12 months ago

Post

4941

InSPyReNet Background Removal

I've built a space for fast background removal.

- gokaygokay/Inspyrenet-Rembg

- https://github.com/plemeri/InSPyReNet

2 replies

reacted to gokaygokay's post with 👍 about 1 year ago

Post

3032

I've created a Stable Diffusion 3 (SD3) image generation space for convenience. Now you can:

1. Generate SD3 prompts from images
2. Enhance your text prompts (turn 1-2 words into full SD3 prompts)

https://huggingface.co/spaces/gokaygokay/SD3-with-VLM-and-Prompt-Enhancer

These features are based on my custom models:

- VLM captioner for prompt generation:
- gokaygokay/sd3-long-captioner

- Prompt Enhancers for SD3 Models:
- gokaygokay/Lamini-Prompt-Enchance-Long
- gokaygokay/Lamini-Prompt-Enchance

You can now simplify your SD3 workflow with these tools!

reacted to gokaygokay's post with 👍🤯🔥 about 1 year ago

Post

6207

Kolors with VLM support

I've built a space for using Kolors image generation model with captioner models and prompt enhancers.

- Space with VLM and Prompt Enhancer
gokaygokay/KolorsPlusPlus

- Original Space for model
gokaygokay/Kolors

- Captioner VLMs
- gokaygokay/sd3-long-captioner-v2

- microsoft/Florence-2-base

- Prompt Enhancers
- gokaygokay/Lamini-Prompt-Enchance-Long

- gokaygokay/Lamini-Prompt-Enchance

reacted to sanchit-gandhi's post with 👀 over 1 year ago

Post

Why does returning timestamps help Whisper reduce hallucinations? 🧐

Empirically, most practitioners have found that setting return_timestamps=True helps reduce hallucinations, particularly when doing long-form evaluation with Transformers’ “chunked” algorithm.

But why does this work?..

My interpretation is that forcing the model to predict timestamps is contradictory to hallucinations. Suppose you have the transcription:

The cat sat on the on the on the mat.

Where we have a repeated hallucination for “on the”. If we ask the model to predict timestamps, then the “on the” has to contribute to the overall segment-level timing, e.g.:

<|0.00|> The cat sat on the on the on the mat.<|5.02|>

However, it’s impossible to fit 3 copies of “on the” within the time allocation given to the segment, so the probability for this hallucinatory sequence becomes lower, and the model actually predicts the correct transcription with highest probability:

<|0.00|> The cat sat on the mat.<|5.02|>

In this sense, the end timestamp is of the opposite of the initial timestamp constraint they describe in Section 4.5 of the paper Robust Speech Recognition via Large-Scale Weak Supervision (2212.04356) → it helps the model remove extra words at the end of the sequence (rather than the initial timestamp which helps when the model ignores words at the start), but the overall principle is the same (using timestamps to improve the probability of more realistic sequences).

Leaving it open to you: why do you think timestamps reduces Whisper hallucinations?