OpenGVLab/Ask-Anything — 3.3k★ trên GitHub (Python). [CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
Tóm tắt dựng từ metadata GitHub của chính dự án — chưa có bài review TopGit. Trang sẽ tự động cập nhật khi bài review đầy đủ được xuất bản.
VÌ SAO CHƯA CÓ REVIEW
TopGit viết bài đầy đủ cho repo có nhiều sao nhất và được yêu cầu nhiều nhất. Trang này là snapshot trong thời gian chờ — xem README gốc ở tab READ ME.
|
|
|
|
[VideoChat-7B-8Bit] End2End ChatBOT for video and image. [InternVideo2-Chat-8B-HD]
中文 README 及 中文交流群 | Paper
⭐️: We are also working on a updated version, stay tuned!
:fire: Updates
2026/07/17: 🚀🚀 We release VideoChat3, a fully open, efficient 4B Video MLLM for general, long-form, and streaming video understanding. VideoChat3 improves 18/19 offline and 10/11 streaming metrics over Qwen3-VL-4B. We release the model weights, code, training recipes, and complete datasets. Check out our paper and homepage!
2025/01/18: We release videochat-flash and videochat-tpo to extend MLLMs' capabilities on both long and accurate video understanding. videochat-flash sets new records in mutiple video benchmarks (for both short and long videos), improving code usability by leveaging LLaVA and others. videochat-tpo exploits classical vision task annotations (e.g. tracking) to optimize MLLMs in a DPO manner, enhancing MLLMs' performance and enabling capabilities in tracking, segmentation, and more.
2024/06/25: We release the branch of videochat2 using vllm, speed up the inference of videochat2.
2024/06/19: 🎉🎉 Our VideoChat2 achieves the best performances among the open-sourced VideoLLMs on MLVU, a multi-task long video understanding benchmark.
2024/06/13: Fix some bug and give testing scripts/
:warning: We replace some repeated (~30) QAs in MVBench, which may only affect the results by 0.5%.
:loudspeaker: We give the scripts for testing EgoSchema and Video-MME, please check the demo_mistral.ipynb and demo_mistral_hd.ipynb.
2024/06/07: :fire::fire::fire: We release VideoChat2_HD, which is fine-tuned with high-resolution data and is capable of handling more diverse tasks. It showcases better performance on different benchmarks, especially for detailed captioning. Furthermore, it achieves 54.8% on Video-MME, the best score among 7B MLLMs. Have a try! 🏃🏻♀️🏃🏻
2024/06/06: We release VideoChat2_phi3, a faster model with robust performaces.
2024/05/22: We release VideoChat2_mistral, which shows better capacity on diverse tasks (60.4% on MVBench, 78.6% on NExT-QA, 63.8% on STAR, 46.4% on TVQA, 54.4% on EgoSchema-full and 80.5% on IntentQA). More details have been updated in the paper.
2024/04/05 MVBench is selected as Poster (Highlight)!
2024/2/27 MVBench is accepted by CVPR2024.
2023/11/29 VideoChat2 and MVBench are released.
VideoChat2 is a robust baseline built on UMT and Vicuna-v0.
2M diverse instruction data are released for effective tuning.
MVBench is a comprehensive benchmark for video understanding.
2023/05/11 End-to-end VideoChat and its technical report.
VideoChat1: Instruction tuning for video chatting (also supports image one).
Paper: We present how we craft VideoChat with two versions (via text and embed) along with some discussions on its background, applications, and more.
2023/04/25 Watch videos longer than one minute with chatGPT
VideoChat LongVideo: Incorporating langchain and whisper into VideoChat.
2023/04/21 Chat with MOSS
VideoChat with MOSS: Explicit communication with MOSS.
2023/04/20: Chat with StableLM
VideoChat with StableLM: Explicit communication with StableLM.
2023/04/19: Code release & Online Demo
VideoChat with ChatGPT: Explicit communication with ChatGPT. Sensitive with time.
MiniGPT-4 for video: Implicit communication with Vicuna. Not sensitive with time. (Simple extension of MiniGPT-4, which will be improved in the future.)
If you find this project useful in your research, please consider cite:
@article{2023videochat,
title={VideoChat: Chat-Centric Video Understanding},
author={KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao},
journal={arXiv preprint arXiv:2305.06355},
year={2023}
}
@inproceedings{li2024mvbench,
title={Mvbench: A comprehensive multi-modal video understanding benchmark},
author={Li, Kunchang and Wang, Yali and He, Yinan and Li, Yizhuo and Wang, Yi and Liu, Yi and Wang, Zun and Xu, Jilan and Chen, Guo and Luo, Ping and others},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={22195--22206},
year={2024}
}
🌤️ Discussion Group
If you have any questions during the trial, running or deployment, feel free to join our WeChat group discussion! If you have any ideas or suggestions for the project, you are also welcome to join our WeChat group discussion!
We are hiring researchers, engineers and interns in General Vision Group, Shanghai AI Lab. If you are interested in working with us, please contact Yi Wang ([email protected]).
OpenGVLab/Ask-Anything thuộc nhóm AI Tools trên TopGit, cùng 13 topic GitHub. Trang Trending và Topics liệt kê các repo cùng số sao và cùng ngôn ngữ để so sánh.
Đọc thêm về OpenGVLab/Ask-Anything ở đâu?
Trang TopGit này là một snapshot — tab "Readme" hiển thị nguyên văn README của repo (đã bỏ link, giữ ảnh). Repo GitHub ở github.com/OpenGVLab/Ask-Anything là nguồn chính thức.
OpenGVLab/Ask-Anything có bao nhiêu sao?
OpenGVLab/Ask-Anything có 3.3k sao GitHub — tải lại trang để xem số mới nhất, hoặc xem trực tiếp github.com/OpenGVLab/Ask-Anything. TopGit phản chiếu số sao của GitHub nhưng không cam kết đến từng phút.
OpenGVLab/Ask-Anything có những chủ đề gì?
GitHub topics của OpenGVLab/Ask-Anything: "big-model", "captioning-videos", "chat", "chatgpt", "foundation-models", "gradio", "langchain", "large-language-models", "large-model", "stablelm", "video", "video-question-answering", "video-understanding". TopGit xếp repo vào nhóm AI Tools.
OpenGVLab/Ask-Anything còn đang phát triển không?
Commit gần nhất trên OpenGVLab/Ask-Anything là 28 ngày trước (theo timestamp GitHub). Repo có 269 fork — một chỉ báo về mức độ quan tâm của cộng đồng.
OpenGVLab/Ask-Anything viết bằng ngôn ngữ gì?
OpenGVLab/Ask-Anything chủ yếu viết bằng Python. Trường "language" của GitHub dựa trên phần lớn byte ở nhánh mặc định.
Vì sao OpenGVLab/Ask-Anything được xếp vào nhóm AI Tools?
TopGit xếp OpenGVLab/Ask-Anything vào nhóm AI Tools dựa trên GitHub topics và mô tả của repo (gắn thẻ: "big-model", "captioning-videos", "chat"). Việc phân loại dựa trên metadata thật của repo, không phải đoán theo cảm tính biên tập.
Đọc đầy đủ README ở tab phía trên.
Vẫn đang phân vân về Ask-Anything?
Một cú bấm sẽ gửi câu hỏi kèm trang này cho AI — xem AI nói gì về Ask-Anything.