Start typing to search this publication.
MetaEnd logo MetaEnd
Open menu
MetaEnd logo

Subscribe to MetaEnd

Get new posts delivered straight to your inbox.

MetaEnd — "MetaEnd" delves into the frontier of AI and blockchain through in-depth discussions on innovative tools, coding techniques, and their multifaceted impact, complemented by daily industry news updates.


A private two-model AI stack on one 8GB GPU

Some data should never leave the building. Bank statements, payment rails, customer records, anything you would not paste into a hosted chatbot. We wanted an assistant that can reason over exactly that kind of material and call real tools against it, with a simple rule: no third party ever sees the prompt or the data. The only network traffic the models make is to localhost. The constraint that makes this interesting is the hardware. One consumer GPU with 8GB of VRAM (an RTX 3070 Ti), a six-c...

Cover image for A private two-model AI stack on one 8GB GPU

Subscribe to MetaEnd

"MetaEnd" delves into the frontier of AI and blockchain through in-depth discussions on innovative tools, coding techniques, and their multifaceted impact, complemented by daily industry news updates.