Science Memes

15572 readers

2604 users here now

Welcome to c/science_memes @ Mander.xyz!

A place for majestic STEMLORD peacocking, as well as memes about the realities of working in a lab.

Rules

Don't throw mud. Behave like an intellectual and remember the human.
Keep it rooted (on topic).
No spam.
Infographics welcome, get schooled.

This is a science community. We use the Dawkins definition of meme.

Research Committee

[email protected]

Other Mander Communities

Science and Research

Biology and Life Sciences

Physical Sciences

Humanities and Social Sciences

Practical and Applied Sciences

Memes

Miscellaneous

founded 2 years ago

MODERATORS

[email protected]

387

Sardonic Grin (mander.xyz)

submitted 1 year ago* (last edited 1 year ago) by [email protected] to c/[email protected]

66 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] [email protected] 7 points 1 year ago (1 children)

This actually is a symptom from the sort of "beneficial" overfit in Deep Learning. As someone whose research is in low data, long tails, and few shot learning, there's a few things that smaller networks did better in generalization, and one thing they particularly did better (without explicit training for it) is gauging uncertainty. This uncertainty is sometimes referred to as calibration. Calibrating deep networks can yield decent probabilities that can be used to show uncertainty.

There are other tricks for this. My favorite strategies prep the network for learning new things. Large margin training and the like are a good thing to look into. Having space in the output semantic space (the layer immediately before the output or earlier for encoder decoder style networks) allows for larger regions for distinct unknown values to be separated from the known ones, which helps inherently calibrate the network.

[+] [email protected] 1 points 1 year ago (1 children)

[removed by mod]

[–] [email protected] 8 points 1 year ago (1 children)

That's LLM bull. The model already knows hangman; it's in the training data. It can introduce variations on the data, especially in response to your stimuli, but it doesn't reinvent that way. If you want to see how it can go astray ask it about stuff you know very well, and watch how it's responses devolve. Better yet, gaslight it. It's very easy to convince LLMs that they're wrong because they're usually trained for yes-manning and non confrontation.

Now don't get me wrong, LLMs are wicked neat, but they don't come up with new ideas, but they can be pushed towards new concepts, even when they don't grasp them. They're really good at sounding sure of themselves, and can easily get people to "learn" new "facts" from them, even when completely wrong. Always look up their sources, (which Bard (Google's) can natively get for you in its UI) but enjoy their new ideas for the sake of inspiration. They're neat toys, which can be used to provide natural language interfaces to expert systems. They aren't expert systems.

But also, and more importantly, that's not zero-shot learning. Neat little anecdote from a conversation with them though. Which model are you using?