1. Phonon-2

    Fermion Research's open English speech-to-text model: Parakeet-level accuracy in a 164 MB file, 174 times realtime on a MacBook Air, CC-BY-4.0.
  2. AuK

    Tencent's open-source 1.5B speech model: text to speech, voice cloning, and editing of recorded speech (words, pitch, speed, emotion, noise) from plain-language instructions.
  3. Recordly

    Free AGPL screen recorder with auto-zoom, cursor polish and a timeline editor. The open-source answer to Screen Studio, on Mac, Windows and Linux.
  4. HyperFrames

    HeyGen's open-source framework for rendering HTML/CSS/JS compositions to MP4, MOV, or WebM from the terminal, with agent-first authoring built in.