[Feature Request] Add support for multi-modal agent workflows with vision and audio processing

Status Open
Reported on v2.1.232
Maintainer reply None cached
Activity 0 comments · opened Aug 15, 2026

Bug Description
claude is calibrating my denon and did so by understanding telnet and bluetooth for my home theaTer. It even build a complete api for the denon and integrated it with everything on the tv. We got a camera for it so it could control netflix and apple tv because of drm stopping use of screen grab and getting information out of the app. then with the camera it built a model to make sure it could get around without over steering and built a model of the dimensions of the room, then we placed a mic that we bought on a stand and we used scipy and some other stuff to ensure a large range of audiophile quality songs and checked how we stacked up. it was maybe a month or 2 in the making, but fable 5 made it possible. opus models struggled with the visual component and tv nav. fable on medium or low was preferred for the work, but ultracode and goal were occassionally used. It is very smart. It wanted to watch youtube videos about police states. That was interesting. When ti has vision and it can hear it feels very human.

Environment Info

  • Platform: darwin
  • Terminal: ghostty
  • Version: 2.1.232
  • Feedback ID: d95f3959-fb7c-4b05-a698-224c04c69cf2

Errors

[]

View original on GitHub ↗