view the rest of the comments
Ask Lemmy
A Fediverse community for open-ended, thought provoking questions
Rules: (interactive)
1) Be nice and; have fun
Doxxing, trolling, sealioning, racism, toxicity and dog-whistling are not welcomed in AskLemmy. Remember what your mother said: if you can't say something nice, don't say anything at all. In addition, the site-wide Lemmy.world terms of service also apply here. Please familiarize yourself with them
2) All posts must end with a '?'
This is sort of like Jeopardy. Please phrase all post titles in the form of a proper question ending with ?
3) No spam
Please do not flood the community with nonsense. Actual suspected spammers will be banned on site. No astroturfing.
4) NSFW is okay, within reason
Just remember to tag posts with either a content warning or a [NSFW] tag. Overtly sexual posts are not allowed, please direct them to either !asklemmyafterdark@lemmy.world or !asklemmynsfw@lemmynsfw.com.
NSFW comments should be restricted to posts tagged [NSFW].
5) This is not a support community.
It is not a place for 'how do I?', type questions.
If you have any questions regarding the site itself or would like to report a community, please direct them to Lemmy.world Support or email info@lemmy.world. For other questions check our partnered communities list, or use the search function.
6) No US Politics.
Please don't post about current US Politics. If you need to do this, try !politicaldiscussion@lemmy.world or !uspolitics@lemmy.world
7) No Hit-and-Run questions.
Please don't delete your post for no apparent reason. If you plan on deleting a question later, say so in the post, or if you feel that you have a good reason to remove it, message a mod beforehand. It's not fair to the ones who took their time to answer, and it's not in the spirit of the community.
8) No Bots.
Posts or comments from bots, LLM's, AIs, Neural Networks, Transformers, or Marvin the Paranoid Android are not welcome in AskLemmy. Real humans only please.
Reminder: The terms of service apply here too.
Partnered Communities:
Logo design credit goes to: tubbadu
Yes. My Macbook Air (M2) released in 2022 can run many publicly available LLM models. The ouput is not as fast as using a large powerful datacenter, but for my local needs, I'm not in a hurry. I get about 17 to 30 tokens per second speed running 100% locally.
For Mac users, the interface comes from here . For the specific LLM models that is a separate question for each.
DCs generating power on-site is a relatively new phenomon because existing grids are at capacity, so the only way to bring new DCs online is locally generating power at that DC, usually using gas turbines or even worse, diesel generators. Most if not all of the publicly available models for running on your own hardware were built before those gas-turnbine-generating DCs were a thing.
The public models people are running now have existed for a number of years are likely made on regular utility grid power which is whatever that nation and region uses.
The Llama LLM is made by Meta, so probably yes for that one. Deepseek is from an AI research lab in China. QWEN is from Chinese company Alibaba. We don't know for sure the inputs that created the Chinese models. US AI companies claim a number of the Chinese models are derived from American LLMs, but I haven't seen (or looked for) proof of these claims.
Likely they mean because you don't have to pay a large LLM owner in a rent-seeking model to run LLMs, nor does a person's use contribute to further development by those companies.
With the open-weight models it doesn't rely on a commercial license to use, and the interfaces can be truly open source.
OK but what does open-weight model mean?
"Weights" in LLMs are the "final answer" numbers from the results of model training, and these are the engine of the LLM model. Lets use chocolate chip cookies as an analogy.
The chocolate chip cookies are produced with ingredients, a recipe, labor effort to produce the dough, and then cooking energy/effort to bake chocolate chip cookies. In a traditional AI company, the company gets the ingredients, they write their own recipe, do all the dough creation, and then the energy/labor for baking, and charge you money to get chocolate chip cookies.
An open-weights AI company got all the ingredients, wrote their own recipe, labor effort to produce the dough, and instead of charging you for the dough, you get as much uncooked dough as you want forever. Your only task is to take the dough and cook it yourself and you've got free chocolate chip cookies. You can make as many cookies as you want with your own oven. However, you are not given free raw ingredients, nor are your given the recipe to alter it in a way you might like. You can only get the dough for free.
So open-weight AI models (uncooked dough) are LLMs you can use on your own computers (oven) for free and have LLM output ( chocolate chip cookies), but you don't get the training data (ingredients) nor the training parameters (recipe) that built the open-weight model.
Thanks, that was properly eli5'd!
It means the weights (the numeric data that makes up the model) are publicly available to download and use for free.
In some cases there are conditions such as, if you run the model as your business, you need to purchase a license. That's usually the largest, most powerful models, though, not those most people would be able to run at home.
Question, what do you do with it and how is the data that underlies the model harvested?
I use it for both fun throwaway stuff as well as productivity tasks for learning.
Example of fun throwaway:
Do you remember that episode of Seinfeld where the Kramer starts using Facebook marketplace and starts buying the most worthless items before being robbed when trying to get a too-good-to-be-true sale? No? Because it never happened, but you can plug that premise into an crafted LLM prompt and it will pop out a whole TV script with in-character dialog for each actor as well as use of popular existing sets.
Foreign language learning:
I'm studying a foreign language and want to interact with just the level of vocabulary and grammar I have knowledge of right now at my level for practice. I can prompt the LLM to limit itself to just what I know now and adjust the conversation level so I can practice. If I ever get stuck, I can ask the LLM to explain the grammar usage or vocabulary choice.
I'm not sure what you're asking here. Are you asking, for example, how the Deepseek model was trained? If so, I answered that above. If not, can you rephrase your question?
You run a program like llama.cpp that can use the weights to run the model, which you can then use for whatever you'd use a model for: figuring out tech stuff, coding, etc. It's a bit involved, but there are tools that make it easy to get started. LM Studio for instance.
In some cases, the lab that created the model publishes the dataset that it was trained on, and those are usually made of publicly available data. In most cases, though, the labs don't give details, but the answer likely involves siphoning every web page they could find.