Yes, we are.
I mean the latter statement is not true at all. I'm not sure why you think this. A basic GPT model reads a sequence of tokens and predicts the next one. Any sequence of tokens is possible, and each digit 0-9 is likely its own token, as is the case in the GPT2 tokenizer.
An LLM can't generate random numbers in the sense of a proper PRNG simulating draws from a uniform distribution, the output will probably have some kind of statistical bias. But it doesn't have to produce sequences contained in the training data.
I'm ok with accepting this as canon
It's called speed of lobsters
If you fine tune a LLM on math equations, odds are it won't actually learn how to reliably solve novel problems. Just the same as it won't become a subject matter expert on any topic, but it's a lot harder to write simple math that "looks, but is not, correct" than it is to waffle vaguely about a topic. The idea of a LLM creating a robust model of the semantics of the text it's trained on is, at face value, plausible; it just doesn't seem to actually happen in practice.
If you have a lot of semantic breakpoints (like the end of a concept) that don't line up with syntactic breakpoints (like the end of a method or expression body) your code probably needs to be refactored. If you don't, then automatic code formatting is probably all you need.
Do I work with you
True, but that's definitely C#
Outdoors
I'm a professional software developer with ML experience, albeit not an expert in ML specifically. It would obviously affect the literal value of the embeddings, but there's no chance it would have a qualitative effect on a reasonably performant model.
Syntax highlighting, linting, and language specific autocomplete are features supported by plugins and scripts. Plain, simple vim is a powerful extensible text editor. The extensibility makes it easy to turn into an IDE.
kogasa
0 post score0 comment score
Probably a good thing. Some kind of anxiety where you know you have to say something so you practice in your head over and over, and then replay it afterwards over and over just thinking if it sounded weird or if you could have done it better.