Johann-Peter Hartmann PRO
johannhartmann
AI & ML interests
LLMs, Local LLMs, Transformers, Image Processing, Audio Processing, E-Commerce
Recent Activity
liked a model 9 days ago
Soofi-Project/Soofi-S-Isar-Preview upvoted a collection 9 days ago
Soofi S Beta Models liked a Space 10 days ago
Soofi-Project/Pretraining-Tech-Report-oldOrganizations
Document & UI Intelligence
-
xlangai/Aguvis-7B-720P
8B • Updated • 28 • 9 -
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Paper • 2412.04454 • Published • 70 -
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Paper • 2401.10935 • Published • 6 -
cckevinn/SeeClick
Text Generation • 10B • Updated • 391 • 18
Medical MultiModal
Multimodal models that have been trained on medical datasets.
Computer Use Models
-
ByteDance-Seed/UI-TARS-72B-DPO
Image-Text-to-Text • 73B • Updated • 1.53k • 157 -
ByteDance-Seed/UI-TARS-7B-DPO
Image-Text-to-Text • 8B • Updated • 1.83k • 228 -
microsoft/OmniParser
Image-Text-to-Text • Updated • 353 • 1.71k -
jadechoghari/Ferret-UI-Llama8b
Image-Text-to-Text • 8B • Updated • 440 • 68
Multimodal Models
A collection of multimodal models for the gpu poor
-
google/paligemma-3b-pt-896
Image-Text-to-Text • 3B • Updated • 761 • 125 -
OpenGVLab/InternVL-Chat-V1-5
Image-Text-to-Text • 26B • Updated • 6.18k • 417 -
alexshengzhili/llava-v1.5-13b-dpo
Text Generation • Updated • 8 • 5 -
llava-hf/llava-v1.6-mistral-7b-hf
Image-Text-to-Text • 8B • Updated • 607k • 312
Music
Computer Use Models
-
ByteDance-Seed/UI-TARS-72B-DPO
Image-Text-to-Text • 73B • Updated • 1.53k • 157 -
ByteDance-Seed/UI-TARS-7B-DPO
Image-Text-to-Text • 8B • Updated • 1.83k • 228 -
microsoft/OmniParser
Image-Text-to-Text • Updated • 353 • 1.71k -
jadechoghari/Ferret-UI-Llama8b
Image-Text-to-Text • 8B • Updated • 440 • 68
Document & UI Intelligence
-
xlangai/Aguvis-7B-720P
8B • Updated • 28 • 9 -
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Paper • 2412.04454 • Published • 70 -
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Paper • 2401.10935 • Published • 6 -
cckevinn/SeeClick
Text Generation • 10B • Updated • 391 • 18
Multimodal Models
A collection of multimodal models for the gpu poor
-
google/paligemma-3b-pt-896
Image-Text-to-Text • 3B • Updated • 761 • 125 -
OpenGVLab/InternVL-Chat-V1-5
Image-Text-to-Text • 26B • Updated • 6.18k • 417 -
alexshengzhili/llava-v1.5-13b-dpo
Text Generation • Updated • 8 • 5 -
llava-hf/llava-v1.6-mistral-7b-hf
Image-Text-to-Text • 8B • Updated • 607k • 312
Medical MultiModal
Multimodal models that have been trained on medical datasets.