www.nature.com
Evelyn Boyd Granville, space-flight trailblazer (1924—2023)
Mathematician and programmer who transcended barriers of race and gender. Mathematician and programmer who transcended barriers of race and gender.
#STS is an active hashtag on Bluesky. In the last 30 days, 62 people shared 126 posts with it — around 4 a day. Activity is up 0% versus the previous week, peaking on Sep 21 with 8 posts.
Tags most often used together with #STS.
www.nature.com
Evelyn Boyd Granville, space-flight trailblazer (1924—2023)
Mathematician and programmer who transcended barriers of race and gender. Mathematician and programmer who transcended barriers of race and gender.
www.oumnh.ox.ac.uk
Darwin's Queer Theory
Queer ecology is a vibrant and rapidly evolving field, exploring the astonishing diversity of sexes and sexual repertoires found among plants and animals. For some, this has been a revelation.
www.scientificamerican.com
Removing Race from Lung Function Tests Could Benefit Millions of Black Americans
A new study shows that hundreds of thousands more Black people in the U.S. would qualify for a lung disease diagnosis and disability payments if lung-function measurements weren’t adjusted for race
This week, the UniMelb HPS Seminar was lucky to hear PhD candidate & founding podcast host Samara Greenwood deliver a wonderful completion seminar: 'Working through Context: A Study of Contextual Explanation in HPS' Join us in wishing Samara the best for her final push towards completing her PhD! 🎉
www.4sonline.org
About the Conference
www.theguardian.com
Chicago Sun-Times confirms AI was used to create reading list of books that don’t exist
Outlet calls story, created by freelancer working with one of the newpaper’s content partner, a ‘learning moment’
www.oumnh.ox.ac.uk
Darwin's Queer Theory
Queer ecology is a vibrant and rapidly evolving field, exploring the astonishing diversity of sexes and sexual repertoires found among plants and animals. For some, this has been a revelation.
bit.ly
Improving reliability of large language models via claim-level self-verification and uncertainty calibration - Discover Artificial Intelligence
Large Language Models (LLMs) often generate fluent answers that appear confident even when they contain factual errors. This creates a reliability problem because users may trust incorrect answers when the model provides no meaningful signal of uncertainty. This paper proposes CLAIM-CAL, a claim-level self-verification framework for calibrated LLM reliability. Instead of assigning confidence to a complete answer directly, CLAIM-CAL decomposes an answer into atomic factual claims, verifies each claim using multiple verification probes, converts claim-level verdicts into a risk score, and then applies post-hoc calibration to obtain a more reliable confidence estimate. We evaluate CLAIM-CAL on the TruthfulQA generation benchmark using a 200-example experimental run with 60 examples for calibration and 140 examples for held-out testing. The proposed method is compared against direct answering, verbal confidence, self-consistency, and simple self-verification. In the held-out comparison, raw CLAIM-CAL achieved the highest observed accuracy among the evaluated methods but remained overconfident. Isotonic calibration reduced Expected Calibration Error (ECE), a bin-weighted gap between predicted confidence and empirical accuracy, from 0.212 for raw CLAIM-CAL to 0.038 [95% CI: 0.015, 0.106] after calibration, while maintaining accuracy at 0.757. To address calibration as a possible confound, we additionally applied isotonic calibration to the variable-confidence baselines on the same 60-example calibration split. In this fair calibrated comparison, CLAIM-CAL + Calibration achieved the strongest point estimate for 10-bin equal-width ECE (0.038) and Brier Score (0.146), although paired bootstrap confidence intervals show that ECE differences versus the closest calibrated baselines should be interpreted cautiously. The calibrated method also achieved selective accuracy of 0.824 at a 0.7 confidence threshold with 0.893 coverage. A second-pass judge prompt-robustness check on 50 sampled evaluations achieved 96.0% agreement and Cohen's kappa of 0.896. These findings suggest that claim-level verification produces a useful reliability signal, and that calibration is necessary to convert that signal into trustworthy confidence estimates for selective answering.
This week, the UniMelb HPS Seminar was lucky to hear PhD candidate & founding podcast host Samara Greenwood deliver a wonderful completion seminar: 'Working through Context: A Study of Contextual Explanation in HPS' Join us in wishing Samara the best for her final push towards completing her PhD! 🎉
Posts are pulled live from Bluesky and cached briefly. Posts with content labels are hidden.