Skip to content

Assessment Pipeline: Add audio/video support #1184

Description

@kartpop

Is your feature request related to a problem?
The AI Assessments pipeline currently lacks support for audio and video modalities. This limits the assessment's effectiveness and prevents a comprehensive analysis of inputs.

Describe the solution you'd like

  • Extend multimodal support to include audio/video in AI Assessments.
  • Experiment with Gemini for video handling and explore methods with OpenAI/Anthropic.
  • Sample an audio/video dataset from partners, run the pipeline, conduct human evaluations, and iterate.
  • Support a language mix of ~50–60% English, ~10–15% Tamil, and a strong presence of South Indian languages like Telugu.
Original issue

Context

Extend multimodal support to include audio/video in AI Assessments.

Investigation

  • Video: Gemini does appear to support video directly. Experiments to be done to figure out reasonable way of handling video in Open AI / Anthropic (sampling frames + full audio transcript ?)

Approach / acceptance criteria

  • Sample audio/video dataset from partners, run pipeline, run human evals, iterate. Building the tech is easy; evaluation is the bottleneck.
  • Language mix to support: ~50–60%+ English, Tamil ~10–15%, strong South Indian language presence, Telugu

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions