Post
2193
Demo of OCR & Math QA using multi-capable VLMs like MonkeyOCR-pro-1.2B, R1-One-Vision, VisionaryR1, Vision Matters-7B, and VIGAL-7B, all running together with support for both image and video inference. πͺ
β¦ Demo Spaces :
β€· Multimodal VLMs : prithivMLmods/Multimodal-VLMs
β¦ Models :
β€· Visionary R1 : maifoundations/Visionary-R1
β€· MonkeyOCR [1.2B] : echo840/MonkeyOCR-pro-1.2B
β€· ViGaL 7B : yunfeixie/ViGaL-7B
β€· Lh41-1042-Magellanic-7B-0711 : prithivMLmods/Lh41-1042-Magellanic-7B-0711
β€· Vision Matters 7B : Yuting6/Vision-Matters-7B
β€· WR30a-Deep-7B-0711 : prithivMLmods/WR30a-Deep-7B-0711
β¦ MonkeyOCR-pro-1.2B Colab T4 Demo [ notebook ]
β€· MonkeyOCR-pro-1.2B-ReportLab : https://github.com/PRITHIVSAKTHIUR/OCR-ReportLab/blob/main/MonkeyOCR-0709/MonkeyOCR-pro-1.2B-ReportLab.ipynb
β¦ GitHub : https://github.com/PRITHIVSAKTHIUR/OCR-ReportLab
The community GPU grant was given by Hugging Face β special thanks to them.π€π
.
.
.
To know more about it, visit the model card of the respective model. !!
β¦ Demo Spaces :
β€· Multimodal VLMs : prithivMLmods/Multimodal-VLMs
β¦ Models :
β€· Visionary R1 : maifoundations/Visionary-R1
β€· MonkeyOCR [1.2B] : echo840/MonkeyOCR-pro-1.2B
β€· ViGaL 7B : yunfeixie/ViGaL-7B
β€· Lh41-1042-Magellanic-7B-0711 : prithivMLmods/Lh41-1042-Magellanic-7B-0711
β€· Vision Matters 7B : Yuting6/Vision-Matters-7B
β€· WR30a-Deep-7B-0711 : prithivMLmods/WR30a-Deep-7B-0711
β¦ MonkeyOCR-pro-1.2B Colab T4 Demo [ notebook ]
β€· MonkeyOCR-pro-1.2B-ReportLab : https://github.com/PRITHIVSAKTHIUR/OCR-ReportLab/blob/main/MonkeyOCR-0709/MonkeyOCR-pro-1.2B-ReportLab.ipynb
β¦ GitHub : https://github.com/PRITHIVSAKTHIUR/OCR-ReportLab
The community GPU grant was given by Hugging Face β special thanks to them.π€π
.
.
.
To know more about it, visit the model card of the respective model. !!