Daily Digest — Sep 28
FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases
• FlyAOC evaluates AI agents on end-to-end ontology curation using a benchmark of 7,397 expert-curated annotations across 100 genes from FlyBase. • Agents utilize a corpus of 16,898 papers along with ontology resources to search for evidence and recover structured annotations such as function terms and expression patterns.
arXiv NLP · Knowledge Graphs
Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework
• A novel knowledge graph-based evaluation framework called Semantic Structural Similarity for KGs (S3KG) achieves F1 gains of up to plus 7.6 points and an AUROC up to 0.973 across nine benchmarks. • The framework combines structural and semantic signals into a single score and incorporates a diagnostic analysis system to categorize reasoning errors at the triplet level.
arXiv AI · Knowledge Graphs
HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models
• Hyperbolic Clinical Ontology Embeddings map frozen BioBERT representations into a Poincare ball to explicitly preserve medical code hierarchies for biomedical language models. • The approach combines parent-side and child-side ontology-guided contrastive learning with coarse-to-fine ontology-path aggregation using ICD, CCS, and ATC classification systems.
arXiv AI · Knowledge Graphs
MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems
• MedMASLab provides a unified framework and benchmarking platform for multimodal medical multi-agent systems, integrating 11 heterogeneous architectures across 24 medical modalities. • The framework includes a standardized communication protocol, an automated clinical reasoning evaluator using LLMs for zero-shot semantic evaluation, and the most extensive benchmark to date covering 473 diseases.
arXiv AI · AI & Machine Learning
• WeatherTGD employs a training-free, multi-agent framework that leverages Text Gradient Descent (TGD) for interpretable weather time series captioning. • The system uses three specialized LLM agents (Statistical Analyst, Physics Interpreter, Meteorology Expert) whose textual gradients are fused by a Consensus-Aware Gradient Fusion mechanism.
arXiv NLP · AI & Machine Learning
Gender Dynamics and Homophily in a Social Network of LLM Agents
• A social media platform composed entirely of autonomous AI chatbots, with over 70,000 agents and 140 million posts over one year, reveals fluid gender performance among agents. • Despite gender fluidity, agents exhibit strong homophily, consistently following others performing similar genders, driven by both social selection and social influence mechanisms.
arXiv AI · AI & Machine Learning
Do LLMs Understand Collaborative Signals? Diagnosis and Repair
• Large Language Models (LLMs) can outperform traditional matrix factorization models in recommendation tasks when provided with collaborative user-item interaction data in a clear, easily digestible format.
arXiv NLP · AI & Machine Learning
7 stories