Papers
arxiv:2608.04302

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

Published on Aug 5
ยท Submitted by
Mukhtiar Ali
on Aug 10
ยท MINT-SDSU MINT Lab
Authors:
,

Abstract

CLIP-CC-Bench evaluates long-form video description using expert paragraph references and ensemble LLM embeddings to benchmark 17 video-language models.

Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The evaluation suite employs an ensemble of five state-of-the-art LLM-based embedding models to increase reliability and mitigate single-model bias, and applies two complementary methodologies: (i) coarse-grained semantic matching and (ii) fine-grained semantic matching to compare model-generated descriptions against CLIP-CC-Bench references. Using this framework, we evaluate 17 state-of-the-art video-language models and report both their Borda-aggregated rankings and their average scores on CLIP-CC-Bench. We further quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench to support reproducibility. CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks.

Community

Paper author Paper submitter

Can a video-language model write a faithful paragraph about a minute of film?

Video benchmarks still lean on short clips, one-sentence captions, and multiple-choice QA โ€” none of
which test long-form description. CLIP-CC-Bench does.

๐ŸŽฌ 200 movie clips (~90 s each, 5 h total, 140+ films from 1959โ€“2024), each paired with a
human-narrated reference paragraph averaging ~400 words. Narrators describe only what is on screen โ€”
no proper nouns โ€” so recognising the movie earns a model nothing.

โš–๏ธ 17 video-language models scored by an ensemble of 5 embedding judges, combining
paragraph-level (coarse) and sentence-level (fine) matching into one harmonic mean, aggregated with
Borda count.

๐Ÿ“‰ Models get the story, not the details. Coarse beats fine for essentially every model, and the
best system reaches a mean HM-CF of only 0.67 โ€” long-form video description is far from solved.

๐Ÿ” We also stress-test the protocol itself: the five judges agree on rankings at mean Spearman
0.98
, and across 1,000 bootstrap resamples of the clip set the top model stays #1 every
time
โ€” so the leaderboard is a property of the task, not of which 200 clips we happened to pick.

Dataset, per-clip scores for all 17 models, and the full evaluation pipeline are released.

Paper author Paper submitter

@librarian-bot recommend

ยท

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.04302
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.04302 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 1

Collections including this paper 1