Skip to content

Multimodality - #16

Draft
schifyor wants to merge 48 commits into
ad-freiburg:mainfrom
schifyor:multimodality
Draft

Multimodality#16
schifyor wants to merge 48 commits into
ad-freiburg:mainfrom
schifyor:multimodality

Conversation

@schifyor

Copy link
Copy Markdown

Expanding GRASP's multimodal capabilities as part of a Bachelors Project

New features:

  • End-to-end multimodal support: GRASP now handles text + images + audio across backend, reasoning tools (load()/analyze()).
  • Full user input integration: Added multimodal inputs in both CLI (--image-input, --audio-input, --load-user-input) and web UI (multi-image/audio upload + PDF-to-image page selection).
  • Modality-aware model architecture: Refactored from single-model setup to multi-model configuration by modality (GRASP/text/image/audio), with updated OpenAI adapter and new multimodal dependencies/benchmark coverage.

@schifyor
schifyor force-pushed the multimodality branch 2 times, most recently from 76d0b86 to c866d52 Compare July 15, 2026 08:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants