add-image-vision

Solid

Add image vision to Deus agents. Resizes and processes WhatsApp image attachments, then sends them to Claude as multimodal content blocks.

AI & Automation 48 stars 3 forks Updated today MIT

Install

View on GitHub

Quality Score: 81/100

Stars 20%
56
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# Image Vision Skill Adds the ability for Deus agents to see and understand images sent via WhatsApp. Images are downloaded, resized with sharp, saved to the group workspace, and passed to the agent as base64-encoded multimodal content blocks. ## Phase 1: Pre-flight 1. Check if `src/image.ts` exists — skip to Phase 3 if already applied 2. Confirm `sharp` is installable (native bindings require build tools) **Prerequisite:** WhatsApp must be installed first (via `/add-whatsapp`). This skill modifies WhatsApp channel files. ## Phase 2: Apply Code Changes Image vision is part of the WhatsApp MCP package in `packages/`. Check if the image module already exists: ```bash test -f src/image.ts && echo "Already present" || echo "Not present" ``` If not present, the WhatsApp MCP package should include image vision support. Ensure the WhatsApp channel is installed and up to date by running `/add-whatsapp`. The following files are involved: - `src/image.ts` (image download, resize via sharp, base64 encoding) - `src/image.test.ts` (8 unit tests) - Image attachment handling in `src/channels/whatsapp.ts` - Image passing to agent in `src/index.ts` and `src/container-runner.ts` - Image content block support in `container/agent-runner/src/index.ts` - `sharp` npm dependency in `package.json` ### Validate code changes ```bash npm install npm run build npx vitest run src/image.test.ts ``` All tests must pass and build must be clean before proceeding. ## Phase 3: Configure 1. Rebuild...

Details

Author
sliamh11
Repository
sliamh11/Deus
Created
4 months ago
Last Updated
today
Language
Python
License
MIT

Integrates with

Similar Skills

Semantically similar based on skill content — not just same category