Skip to content
@VIPL-Audio-Visual-Speech-Understanding

VIPL AVSU

Audio-Visual Speech Understanding Research Group at Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences, ICT, CAS

Pinned Loading

  1. VIPL-AVSU-Group VIPL-AVSU-Group Public

    Collection of works from VIPL-AVSU

    50 5

  2. LRW1000--CAS-VSR-W1k LRW1000--CAS-VSR-W1k Public

    DenseNet3D Model In "LRW-1000: A Naturally-Distributed Large-Scale Benchmark for Lip Reading in the Wild", https://arxiv.org/abs/1810.06990

    Python 123 19

  3. learn-an-effective-lip-reading-model-without-pains learn-an-effective-lip-reading-model-without-pains Public

    The PyTorch Code and Model In "Learn an Effective Lip Reading Model without Pains", (https://arxiv.org/abs/2011.07557), which reaches the state-of-art performance in LRW-1000 dataset.

    Python 169 36

  4. CAS-VSR-S101 CAS-VSR-S101 Public

    CAS-VSR-S101: A large-scale Mandarin dataset from TV broadcasts for audio-visual speech research

    9 1

  5. CAS-VSR-MOV20 CAS-VSR-MOV20 Public

    CAS-VSR-MOV20: A challenging dataset for Chinese visual speech recognition, consisting of video clips from 20 movies.

    4

  6. CAS-VSR-S68 CAS-VSR-S68 Public

    CAS-VSR-S68: A dataset for lip reading with unseen speakers, spanning 68 hours of news broadcasts.

    7

Repositories

Showing 10 of 12 repositories
  • VIPL-AVSU-Group Public

    Collection of works from VIPL-AVSU

    VIPL-Audio-Visual-Speech-Understanding/VIPL-AVSU-Group's past year of commit activity
    50 5 1 (1 issue needs help) 0 Updated Aug 24, 2026
  • CAS-VSR-MOV20 Public

    CAS-VSR-MOV20: A challenging dataset for Chinese visual speech recognition, consisting of video clips from 20 movies.

    VIPL-Audio-Visual-Speech-Understanding/CAS-VSR-MOV20's past year of commit activity
    4 0 0 0 Updated Mar 13, 2026
  • CAS-VSR-S68 Public

    CAS-VSR-S68: A dataset for lip reading with unseen speakers, spanning 68 hours of news broadcasts.

    VIPL-Audio-Visual-Speech-Understanding/CAS-VSR-S68's past year of commit activity
    7 0 0 0 Updated Mar 13, 2026
  • CAS-VSR-S101 Public

    CAS-VSR-S101: A large-scale Mandarin dataset from TV broadcasts for audio-visual speech research

    VIPL-Audio-Visual-Speech-Understanding/CAS-VSR-S101's past year of commit activity
    9 1 0 0 Updated Mar 13, 2026
  • LRW1000--CAS-VSR-W1k Public

    DenseNet3D Model In "LRW-1000: A Naturally-Distributed Large-Scale Benchmark for Lip Reading in the Wild", https://arxiv.org/abs/1810.06990

    VIPL-Audio-Visual-Speech-Understanding/LRW1000--CAS-VSR-W1k's past year of commit activity
    Python 123 19 1 0 Updated Mar 13, 2026
  • Active-Speaker-Detection Public

    Active speaker detection in videos: a survey

    VIPL-Audio-Visual-Speech-Understanding/Active-Speaker-Detection's past year of commit activity
    0 0 0 0 Updated Mar 9, 2026
  • learn-an-effective-lip-reading-model-without-pains Public

    The PyTorch Code and Model In "Learn an Effective Lip Reading Model without Pains", (https://arxiv.org/abs/2011.07557), which reaches the state-of-art performance in LRW-1000 dataset.

    VIPL-Audio-Visual-Speech-Understanding/learn-an-effective-lip-reading-model-without-pains's past year of commit activity
    Python 169 36 2 1 Updated Sep 12, 2025
  • MAVSR2025-Track1 Public

    Visual Speech Recognition baseline code for MAVSR2025 Track1

    VIPL-Audio-Visual-Speech-Understanding/MAVSR2025-Track1's past year of commit activity
    Python 6 0 0 0 Updated Jun 5, 2025
  • VIPL-Audio-Visual-Speech-Understanding/MAVSR2025-Track2's past year of commit activity
    Python 3 0 1 0 Updated Dec 29, 2024
  • LipNet-PyTorch Public

    The state-of-art PyTorch implementation of the method described in the paper "LipNet: End-to-End Sentence-level Lipreading" (https://arxiv.org/abs/1611.01599)

    VIPL-Audio-Visual-Speech-Understanding/LipNet-PyTorch's past year of commit activity
    Python 238 58 3 0 Updated Sep 21, 2022

Top languages

Loading…

Most used topics

Loading…