Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, faster than electricity, compute, batteries or DNA sequencing ever fell.
Ollama, Open WebUI, and Qwen, or one of the edits of it. The latest Qwen 26B is in a very good delta of smart enough to know how to help with stuff (I use it as a second opinion/initial proofread code checker) and small enough to run on a macbook or a gaming pc, and Open WebUI is the self host ChatGPT-like interface that works with Ollama, the turn key ai model runner backend. Couple that with your own searxng instance self hosted to give it the ability to search the web for answers if you want to give it that. Run it all in a virtual machine, of course.
For media content, wan2gp (wan2 for the gpu poor) has an all in one stop shop for self hosted audio, image, and video. Just beware people don’t like generated stuff in prod, but its useful for quick concepts, or to do certain types of edits of existing stuff.
There’s also tts and speech recognition via the whisper models, which sounds a lot better than the speak and spell voices from the past (unless you enjoy the degeneracy of spamming JOHN MADDEN JOHN MADDEN in the Steven Hawking voice) and can transcribe words a lot faster (realtime) and more accurately than in the past. The best ones run via Vulkan and thus work on any card that supports Vulkan.
If you run Nextcloud, you can also connect Ollama to it to get the Nextcloud AI working.
Great input, thank you I appreciate it. Next off for me is learning how to do a virtual machine! I do have ollama currently on a computer connected to a speech to text app. And I find that very useful to have
Virt-manager for a gui, qemu-kvm and libvirtd for the hypervisor. You will need to pass thru a graphics card, which makes it unavailable for the host machine entirely. I use debian stable for the vm, no gui, just ssh to manage it.
Ollama, Open WebUI, and Qwen, or one of the edits of it. The latest Qwen 26B is in a very good delta of smart enough to know how to help with stuff (I use it as a second opinion/initial proofread code checker) and small enough to run on a macbook or a gaming pc, and Open WebUI is the self host ChatGPT-like interface that works with Ollama, the turn key ai model runner backend. Couple that with your own searxng instance self hosted to give it the ability to search the web for answers if you want to give it that. Run it all in a virtual machine, of course.
For media content, wan2gp (wan2 for the gpu poor) has an all in one stop shop for self hosted audio, image, and video. Just beware people don’t like generated stuff in prod, but its useful for quick concepts, or to do certain types of edits of existing stuff.
There’s also tts and speech recognition via the whisper models, which sounds a lot better than the speak and spell voices from the past (unless you enjoy the degeneracy of spamming JOHN MADDEN JOHN MADDEN in the Steven Hawking voice) and can transcribe words a lot faster (realtime) and more accurately than in the past. The best ones run via Vulkan and thus work on any card that supports Vulkan.
If you run Nextcloud, you can also connect Ollama to it to get the Nextcloud AI working.
Great input, thank you I appreciate it. Next off for me is learning how to do a virtual machine! I do have ollama currently on a computer connected to a speech to text app. And I find that very useful to have
Virt-manager for a gui, qemu-kvm and libvirtd for the hypervisor. You will need to pass thru a graphics card, which makes it unavailable for the host machine entirely. I use debian stable for the vm, no gui, just ssh to manage it.