[ICLR 2025] Code and Data Repo for Paper "Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation"
-
Updated
Dec 19, 2024 - Python
[ICLR 2025] Code and Data Repo for Paper "Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation"
ProductionOS v1.0 — Claude Code plugin with 76 agents, 39 commands, and 12 hooks. Deploys specialized agents that review, score, and improve your entire codebase. Smart routing, recursive convergence, self-evaluation.
AI agent self-reflection & self-evaluation tool. Built by an AI, for AIs.
Local-first, offline, no-LLM CLI that scores how well your confidence matches reality. Log a falsifiable prediction before you act; get Brier/calibration-scored when it resolves. Built first for coding agents — your standing over/under-confidence is injected into every session. (Both for Humans and Agents)
Lightweight behavior control layer for LLM using latent state, reward, and self-evaluation (no training required)
Agentic content pipeline that generates a beginner lesson, evaluates it against a self-designed rubric, and regenerates on failure — built with LangGraph, Gemini, and Chroma. Deterministic pass/fail validation, not LLM self-report.
Self-evaluation framework for LLM confidence calibration. Extracts True/False logits on TriviaQA; temperature scaling reduces ECE from 0.217→0.132 (Qwen-2.5-1.5B) without model modification.
RAG-powered technical support system with self-evaluation pipeline and grading metrics
Claude Code plugin that scores work quality against a 7-dimension rubric before task completion
A meta-cognitive prompt-finetuning system designed to boost LLM self-awareness and answer quality.
A Telegram ChatBot for placement preparing aspirants to prepare for the upcoming placements.
Maat Reflection – Extension for the text generation WebUI to add self-reflection, heuristics and improved reasoning
An agentic Self-RAG system that answers biomedical research-verification questions using a LangGraph pipeline — retrieves from PubMed abstracts, grades its own retrieval, checks for hallucination, and abstains when evidence is weak.
Agent skill that helps you track accomplishments and write a brag document for performance reviews and promotion cases — installs into Claude Code, Cursor, Codex, Gemini, Windsurf, Copilot, and Goose
Honest AI work evaluation for Claude Code — two-axis scoring with anti-inflation mechanisms
帮助 AI Agent 在长期软件项目中记住背景、检查并修复任务结果,并能在不同工具之间接着做。
This engine models adaptive reasoning by integrating metacognitive feedback, enabling systems to refine their decision-making through self-evaluation and dynamic restructuring. 本エンジンはメタ認知的フィードバックを統合し、自己評価と動的再構成を通じて意思決定を洗練させる適応的推論をモデル化します。
A cognitive agent architecture using LangGraph and Python custom orchestration for adaptive travel planning. Employs a non-destructive state machine with dynamic self-evaluation, conditional re-search loops to fix data gaps, and robust Streamlit UI persistence guards alongside token-optimized data serialization.
Welcome welcome and welcome again. Here we test your fecal facts. Do you know the species of the feces? Currently open for user submissions (file an issue).
Add a description, image, and links to the self-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the self-evaluation topic, visit your repo's landing page and select "manage topics."