<span class="vcard">/u/No_Sky9786</span>
/u/No_Sky9786

looking for contributors – trie based memory efficient LLM runner

SALT shrinks a long document down to a fixed size before it is sent to a language model, keeping the sentences that carry the most information. It works with any model, produces a shorter plain-text prompt, and cuts the compute, memory, and wait …