Rather than have one big LLM that can do everything and requires 600GB of VRAM and data center grade power and cooling, why not create a distributed network of specialized LLMs that can work together, in parallel, at scale, to solve problems and accomplish tasks.
The prompt:
Create a diagram that depicts a network of LLMs all intercommunicating with the most efficient form of networked token exchange, co-ordinated by a master LLM. Each LLM should run on it’s own dedicated hardware, and each tailored to a particular function, such as vision-to-text, or text-to-speech, and the various specialized training areas of function LLMs currently demonstrate. The master LLM should have the power and capability to reason and understand the work required, but it does not necessarily need to do the work. It can leverage a network of experts and computational resources in this web of specialized LLMs. The protocol between the LLMs should be fast and efficient, but represent summarized token exchanges, not NVLink level bit movement. The diagram should reflect the similarity of such a network to the human organism and how many sensory components interoperate to perform function. It should also attempt to illustrate the efficiency and scalability of such a network in contrast to the one-big-LLM model.
I gave this to ChatGPT.
It produced this:

I have nothing to add.
Sean Hignett