Bluesky · Hashtag

#MLSky

28
posts · 30d
12
users
1
posts / day
2.3
posts / user
+ 60% vs last week

#MLSky is an active hashtag on Bluesky. In the last 30 days, 12 people shared 28 posts with it — around 1 a day. Activity is up 60% versus the previous week, peaking on Jul 22 with 4 posts.

#MLSky posts per day (last 30 days)

Related tags

Tags most often used together with #MLSky.

Posts with #MLSky

Blake Richards
@tyrellturing.bsky.social
11 months ago
1/4) I’m excited to announce that I have joined the Paradigms of Intelligence team at Google (github.com/paradigms-of...)! Our team, led by @blaiseaguera.bsky.social, is bringing forward the next stage of AI by pushing on some of the assumptions that underpin current ML. #MLSky #AI #neuroscience
Paradigms of Intelligence Team

github.com

Paradigms of Intelligence Team

Advance our understanding of how intelligence evolves to develop new technologies for the benefit of humanity and other sentient life - Paradigms of Intelligence Team

23 11 179
Ted Underwood
@tedunderwood.com
about 1 year ago
New this morning, a Comment I contributed to Nature Computational Science on the interaction between large language models and the humanities. 🧪 🤖 #MLSky rdcu.be/etk07 The link above will be open-access for a month — plus, I'll reply to this post with a link to a permanently open preprint. +

rdcu.be

The impact of language models on the humanities and vice versa

Nature Computational Science - Many humanists are skeptical of language models and concerned about their effects on universities. However, researchers with a background in the humanities are also...

14 55 168
Blake Richards
@tyrellturing.bsky.social
over 1 year ago
1/ Okay, one thing that has been revealed to me from the replies to this is that many people don't know (or refuse to recognize) the following fact: The unts in ANN are actually not a terrible approximation of how real neurons work! A tiny 🧵. 🧠📈 #NeuroAI #MLSky

Why does anyone have any issue with this? I've seen people suggesting it's problematic, that neuroscientists won't like it, and so on. But, I literally don't see why this is problematic...

21 38 152
Ted Underwood
@tedunderwood.com
over 1 year ago
A timely paper exploring ways academics can pretrain larger models than they think, e.g. by trading time against GPU count. Since the title is misleading, let me also say: US academics do not need $100k for this. They used 2,000 GPU hours in this paper; NSF will give you that. #MLSky
$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources

arxiv.org

$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources

Pre-training is notoriously compute-intensive and academic researchers are notoriously under-resourced. It is, therefore, commonly assumed that academics can't pre-train models. In this paper, we seek...

10 12 143
Naomi Saphra
@nsaphra.bsky.social
over 1 year ago
Transformer LMs get pretty far by acting like ngram models, so why do they learn syntax? A new paper by sunnytqin.bsky.social, me, and @dmelis.bsky.social illuminates grammar learning in a whirlwind tour of generalization, grokking, training dynamics, memorization, and random variation. #mlsky #nlp
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization

arxiv.org

Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization

Language models (LMs), like other neural networks, often favor shortcut heuristics based on surface-level patterns. Although LMs behave like n-gram models early in training, they must eventually learn...

5 30 141
Shahab Bakhtiari
@shahabbakht.bsky.social
about 1 year ago
This paper is making the rounds: arxiv.org/abs/2506.21734 A tiny (27M) brain-inspired model trained just on 1000 samples outperforming o3-mini-high on reasoning tasks. #MLSky 🧠🤖
4 25 130
Blake Richards
@tyrellturing.bsky.social
over 1 year ago
This paper looks interesting - it argues that you don’t need adaptive systems like Adam to get good gradient-based training, instead you can just set a learning rate for different groups of units based on initialization: arxiv.org/abs/2412.11768 #MLSky #NeuroAI
No More Adam: Learning Rate Scaling at Initialization is All You Need

arxiv.org

No More Adam: Learning Rate Scaling at Initialization is All You Need

In this work, we question the necessity of adaptive gradient methods for training deep neural networks. SGD-SaI is a simple yet effective enhancement to stochastic gradient descent with momentum (SGDM...

4 12 116
Mark Histed
@markhisted.org
9 months ago
Same for neuroscience. The lack of ability to measure many neurons’ activity, perturb them, and measure intracellular processes and connections is what limits understanding the brain. The key barriers are not algorithms or AI. 🧪#neuroscience 🧠🤖 #MLSky
Anshul Kundaje @anshulkundaje
Francois usually has good takes. But this suggests a bit of cluelessness about what the key barrier to progress in biology is. It's not algorithms or Al. It's still a lack of the ability to measure many important things in cells ie. assay techdev. Perturb-seq is not all u need.
@ François Chollet & @fchollet • 2d
The most powerful scientific instrument of the 21st century isn't the electron microscope or the particle collider. It's the algorithm.
Today, a scientist in biology, physics... Show more
4 18 108
Shahab Bakhtiari
@shahabbakht.bsky.social
over 1 year ago
Interesting paper showing how LLMs change their representational geometry in-context to match a task structure: arxiv.org/abs/2501.00070 Here, the model’s latent representations show a grid structure matching the task. #MLSKy #NeuroAI
1 12 93
Ted Underwood
@tedunderwood.com
almost 2 years ago
New paper devises a metric for feature complexity, and uses it to show that simpler features are learned earlier and tend to be more important. They’re working with ResNet 50 and idk to what extent this applies to diffusion, but: cool map of visual experience. #MLSky arxiv.org/abs/2407.06076
Figure 2: Qualitative Analysis of “Meta-feature” (cluster of features) Complexity.

(Left) A 2D UMAP projection displays the 10,000 features extracted from an image net trained ResNet 50. These features are organized into 150 clusters through K-means clustering, applied to the feature dictionary  D^* . 30 clusters were selected to analyze features at varying complexity levels.

(Right) For each Meta-feature cluster, the average complexity score is computed. This scoring allows classification of features based on their complexity according to the model. Simple features are often similar to color detectors (e.g., grass, sky) or low-frequency patterns (e.g., a bokeh detector) or lines. In contrast, complex features represent parts or structured objects, including shapes resembling ears or curve detectors. Detailed visualizations of individual Meta-features are shown in Appendix B.
3 16 92

Posts are pulled live from Bluesky and cached briefly. Posts with content labels are hidden.