Research dossierORCID 0000-0001-5573-8465

Vaishak Belle

University of Edinburgh · 3 affiliations on record

What this is

Valency read your published work and assembled this dossier automatically. Every figure links to the real paper behind it.

222 works, 2,631 citations, h-index 20, and an explainability survey now cited by people predicting heart disease mortality in Brazil and testing Android apps.

You keep asking the same question in different clothes: what can a system say about what it knows, and can it say it in time to be useful? The only-knowing semantics gave a precise account of what a knowledge base believes. The sum-product network work made that question answerable in polynomial time over mixed discrete and continuous data. The explainability survey asked it of a deployed classifier and a data scientist who has to defend her model. The recent papers on theory of mind in language models ask it of a system that has no knowledge base at all. Reading the corpus from outside, the survey is what most people found first, and it took your name into fields that have nothing to do with logic.

222works in corpus
2010–2026active
Generated August 5, 2026
02
Your research program

Read as one thread, not a list

Over 40 percent of your tagged work sits in artificial intelligence, with logic and machine learning splitting the next third almost evenly at 18 percent each. That split is the shape of the actual research. The only-knowing paper from 2010 and the sum-product network paper from 2018 are doing the same job in different formalisms: pin down exactly what a representation commits to, then make querying that commitment cheap. The 2020 explainability survey is the one outsiders reach for, and it is also the one written for practitioners rather than for the logic literature, which is part of why it travelled. The infinite-domains survey from the same year is the argument that connects them, since it refuses the tidy division where logic gets the discrete world and learning gets the continuous one.

Where your output lives
  • Artificial Intelligencecs.AI43%
  • Logic in Computer Sciencecs.LO18%
  • Machine Learningcs.LG18%
  • Symbolic Computationcs.SC4%
  • Statistical Machine Learningstat.ML4%
Career, by the numbers
222works
2,631citations
20h-index

Your citations are spread rather than stacked: 2,631 across 222 works, with an h-index of 20. The explainability survey is the single biggest draw at 46 citations, and after it the counts fall away fast, with the queerphobia-in-sentiment-analysis paper next at 15.

The papers carrying the load
03
Where your work travels

Where your explainability work landed

The explainability survey is doing most of the travelling, and it has gone somewhere odd. Its heaviest citers are field-defining reviews: a counterfactuals-and-causability survey at 251 citations and a trustworthy-AI overview at 212. Its most recent citers are working scientists who needed a reference for why their model should be interpretable at all, on urban ozone pollution in Chinese cities, mortality from ischemic heart disease in southern Brazil, and Android app testing with deep Q-networks. The nearest neighbours outside computer science follow the same drift. On bioRxiv the closest match is a best-practices workflow for interpretable machine learning in computational biology, and one showing that many models fit neural dynamics equally well while implying different hypotheses. On medRxiv it is class-contrastive explanations rendered for clinicians, Shapley-based antimicrobial stewardship, and a demonstration that saliency maps on chest radiographs can be adversarially steered without changing the prediction. Your tractable-models line travels more narrowly and more precisely: the reference sum-product network survey picks up the hybrid-domain extension, and that is roughly where it stays.

Highest-impact citations
  1. Trustworthy AI: A Computational Perspective

    2022212 citesHaochen Liu, Yaxin Li, et al.

    citation only
  2. To trust or not to trust an explanation: using LEAF to evaluate local linear XAI methods

    2021PeerJ Computer Science83 citesElvio G. Amparore, Alan Perotti, et al.

    Shows LIME and SHAP produce unstable explanations that diverge from their promised theoretical properties, and answers with a metric set and Python framework for judging local linear explanations, exactly the standardisation problem your survey flagged.

    indexed
  3. Towards Accountability in the Use of Artificial Intelligence for Public Administrations

    202163 citesMichele Loi, Matthias Spielkamp

    Maps 16 public-sector AI guideline documents onto a philosophical account of accountability, arguing that delegation to computational systems is imperfect and that auditing needs clearer entitlements.

    indexed
  4. Sum-Product Networks: A Survey.

    2022IEEE Transactions on Pattern Analysis and Machine Intelligence29 citesSanchez-Cauce, Raquel, Paris, Iago, et al.

    The reference survey of sum-product networks, covering definitions, inference and learning algorithms, software and applications, and citing your hybrid-domain extension as part of that record.

    indexed

Closely related preprints, including work that has carried your methods toward the clinic.

Where a paper cites you but we do not hold its full text, we flag it, so you can tell first-hand sources from second-hand mentions.

04
Your collaboration network

Two centers of gravity

Fengxiang He is the one pulling hardest away from you. You share three papers, but his own collaboration record is dominated by Dacheng Tao at 34 shared papers, an order of magnitude more than anything he shares with you, and almost none of his top collaborators appear anywhere in your network. The other direction of divergence is geographic. Ioannis Papantonis, your most frequent collaborator at six papers, now lists Vienna as his current institution, and all six of his papers in view are ones you share. Andreas Bueff has moved to Linköping. Paulius Dilkas is the clearest bridge outward: through him and through Brendan Juba you sit two steps from Kuldeep S. Meel, and through Kwabena Nuamah two steps from Alan Bundy. The shape overall is a hub. The strongest ties run to people whose publication records are largely yours, and the links that leave your circle are carried by a handful of collaborators working on counting and automated reasoning.

Vaishak BelleIoannis Papantonis · 6Weizhi Tang · 6Nijesh Upreti · 5
Clusters
  • Explainability and tractable probabilistic models
  • Neuro-symbolic learning and learning with weak supervision
  • Counting, compilation and logic
  • Language models and question answering
  • Learning theory and trustworthy machine learning
Most frequent collaborators
  1. Ioannis Papantoniscs.AI6papers
  2. Weizhi Tangcs.AI6papers
  3. Nijesh Upretics.AI5papers
  4. Andreas Bueffcs.LG3papers
  5. Paulius Dilkascs.AI3papers
  6. Brendan Juba3papers
  7. Kwabena Nuamah3papers
  8. Fengxiang He3papers
05
People you should meet

Close to your work, not yet in your network

Four people keep turning up next to your recent directions, none at your institutions and none with a paper co-authored with you. Francesco Leofante is closest to your explainability line: he works on counterfactual explanations with formal guarantees and has just published counterfactual scenarios for automated planning, which meets your explanations-as-plans framing from the verification side. YooJung Choi works the theory of tractable probabilistic models, including what they provably cannot approximate. Akihiro Takemura studies reasoning shortcuts, where the logic is satisfied but the concepts are wrong, and has turned finding and repairing them into a constraint satisfaction problem with a verifiable check. Lance Ying builds Bayesian theory-of-mind models that give epistemic language a computational reading. Two of them would push your logic outward into verification and probability, one would meet your weak-supervision semantics with working systems, and one would give the belief work a footing in cognitive modelling.

Francesco Leofante

Established
counterfactual explanations with formal robustness guaranteescounterfactual reasoning in automated planning

He is building formal guarantees around counterfactual explanations, including robustness under model change and behaviour on incomplete inputs, and his counterfactual-scenarios work for automated planning arrives at the same junction your Counterfactual Explanations as Plans paper approaches from the logic side.

No shared institution with you (he is at Imperial College London, with earlier records at Genoa and RWTH Aachen) and no co-authored papers.

YooJung Choi

Established
tractable probabilistic models and probabilistic circuitshardness and robustness of approximate probabilistic inference

Her hardness results say what a tractable model can and cannot approximate, which is the theoretical floor under the query interface your hybrid sum-product network framework builds, and her post-training robustification work asks what a circuit still guarantees once you modify it, close cousin to the limitation your interventions-and-counterfactuals paper identified in circuit transformations.

No shared institution with you (her record lists UCLA, Texas A&M, Michigan State, Arizona State and Seoul National University) and no co-authored papers.

Akihiro Takemura

Rising
reasoning shortcuts formalised as constraint satisfaction in neuro-symbolic learningdifferentiable logic programming that grounds neural outputs to logical atoms

Reasoning shortcuts are the failure mode where a neuro-symbolic system satisfies the logical constraint while getting the concepts wrong, which is precisely the correctness question your neuro-symbolic weak supervision semantics tries to answer. He formalises them as a constraint satisfaction problem, with an ASP-based check for whether a constraint set pins down the intended concept mapping, a repair algorithm, and complexity bounds, and his differentiable logic programming work shows that grounding neural outputs one-to-one onto logical atoms cuts both shortcut types.

No shared institution with you (he is at the National Institute of Informatics in Japan, with earlier records at Osaka, SOKENDAI and York) and no co-authored papers.

Lance Ying

Rising
epistemic language and Bayesian theory of mindsynthesising agent models for grounded theory-of-mind reasoning

He is giving epistemic language a computational semantics through Bayesian theory of mind and synthesising agent models on the fly from language, which is the probabilistic counterpart of the possible-world account your only-knowing work built and the missing half of your ToM-LM approach of delegating belief reasoning to an external executor.

06
Frontiers in your fields, last 90 days

What just landed next to your work

The last three months have been about making logic-plus-learning systems accountable rather than just building more of them. Takemura and Inoue go after reasoning shortcuts, where the logical constraint is satisfied but the concepts underneath are wrong, by putting rules and constraints into one matrix and grounding neural outputs one-to-one onto logical atoms. Bortolotti, Marconato and colleagues come at the same architectures from the other end, wrapping distribution-free coverage guarantees around concept and label predictions at once. On the logic side, epistemic reasoning is being extended with probabilities over actions and with an explicit account of forgetting. Explainability has turned to auditing: confidence guarantees attached to abductive explanations, benchmarks showing explanation quality tracks the data rather than the model, and a legal-reasoning study finding that language models will report solver-backed answers they never actually computed.

07
Where your next paper should go

Five directions, grounded in your work

Each of these takes a formal result you already have and points it at a place where the current literature is working empirically. The circuit and fairness ideas cash in the tractable-inference machinery on questions people are currently answering by search or by discretising the problem away. The only-knowing and epistemic-planning ideas take the belief work out of its home in knowledge representation and into language models and explanation, where the listener's beliefs are treated as a flat model difference. The PAC-semantics idea is the one that would matter most to the neuro-symbolic field right now, since reasoning shortcuts are being patched architecturally with no statement of what the patched system is guaranteed to have learned.

Every source below is real and linked. The directions, and the reading of them, are generated. Check anything you would lean on against your own knowledge of the field.

  1. 01

    Counterfactual explanations computed exactly inside a probabilistic circuit

    Your interventions-and-counterfactuals paper showed that the standard transformations on tractable probabilistic models do not support the causal queries people want from them, and your hybrid sum-product network framework already answers interval-constrained queries over mixed discrete and continuous data. Counterfactual explanation is a causal query that today gets computed by search outside the model, which is why so much of that literature is about robustness after the fact.

    The gap

    Searching the counterfactual-explanations-in-tractable-models neighbourhood surfaces a busy area: gradient-based counterfactuals over tractable probabilistic models from Shao and Kersting, compositional causal inference with tractable circuits from Wang and Kwiatkowska, probabilistic guarantees for robust counterfactuals from Marzari and Leofante, and hardness results from Artelt and colleagues. None computes the counterfactual by exact inference inside the circuit over mixed discrete-continuous features using the weighted model integration machinery your hybrid framework provides.

    First step

    With YooJung Choi, define a counterfactual query over a piecewise-polynomial sum-product network and work out which circuit properties make it answerable in time polynomial in the network size, then check the answer against a search-based counterfactual generator on the same model.

    Bridgescs.LGcs.AIstat.ML
  2. 02

    An only-knowing semantics for auditing what a language model actually commits to

    Only-knowing was invented to say exactly what a knowledge base believes and nothing more, which is a closure condition, not a list of true statements. Benchmarks for belief in language models test whether a model tracks a particular fact, and your own ToM-LM work handed the belief reasoning to an external executor. Nobody has asked the closure question: given a model's answers, what is the maximal set it is committed to, and what does it thereby not believe?

    The gap

    The probe over epistemic logic and language-model belief returns dynamic epistemic logic benchmarks (Sileo and Lernould's MindGames), inference-time scaling with dynamic epistemic logic (Wu and colleagues' DEL-ToM), belief-stability and epistemic-integrity evaluations, mechanistic work on how models track beliefs through lookbacks, and Bayesian epistemology framings. All test what a model asserts. None uses an only-knowing style semantics to characterise the closed set of commitments behind those assertions.

    First step

    With Lance Ying, take the language-augmented Bayesian theory-of-mind setup and pair it with an only-knowing reading of the model's answers, so that a probe set yields both a probabilistic belief estimate and an explicit statement of what the model does not know.

    Bridgescs.AIcs.LOcs.CL
  3. 03

    PAC-semantics guarantees for neuro-symbolic learning under weak supervision

    Your work on implicit learning in first-order logic and with noisy data in linear arithmetic gives validity guarantees without ever writing down the full knowledge base, and your weak-supervision paper gives the semantics for learning from partial labels. Those are the same guarantee in two settings. The neuro-symbolic field is currently discovering the failure modes empirically and patching them architecturally.

    The gap

    The neuro-symbolic guarantees probe returns analyses of the independence assumption (van Krieken and colleagues), hardness of probabilistic neuro-symbolic learning (Maene, Derkinderen and De Raedt), constraint-based and prototypical treatments of reasoning shortcuts, and expectation-maximisation formulations. The guarantees on offer are about optimisation and about hardness. None states a PAC-semantics style validity guarantee for what a neuro-symbolic system has learned when the supervision is partial.

    First step

    With Akihiro Takemura, take a documented reasoning-shortcut benchmark and ask what a PAC-semantics validity statement over the concept predicates would rule out, then test whether that condition is enough to block the shortcut without extra architectural constraints.

    Bridgescs.AIcs.LGcs.LO
  4. 04

    Certifying fairness over continuous attributes with weighted model integration

    Your logical theory of fairness and bias defines the properties in a language that admits real variables, and your knowledge-compilation work scales inference in linear and non-linear hybrid domains. Together those are the two halves of a fairness certificate that does not require discretising income, age or a continuous proxy first. Discretisation is where most current fairness verification quietly loses its guarantee.

    The gap

    The fairness-verification probe returns graphical-model fairness verification (Ghosh, Basu and Meel), quantitative verification for tree ensembles, neural-network fairness checking with Fairify, and probabilistic machine-learning verification via weighted model integration from Morettin, Passerini and Sebastiani. The fairness-specific tools work over discrete or discretised features, and the one weighted-model-integration paper targets general probabilistic properties rather than a logically specified fairness criterion over continuous attributes.

    First step

    Express one non-trivial fairness property from your logical theory as a weighted model integration problem over a classifier with at least one continuous protected proxy, then measure how much the certificate changes when the same attribute is discretised the way current tools require.

    Bridgescs.LOcs.LGcs.CY
  5. 05

    Explanation as a plan over the listener's nested beliefs

    You already read a counterfactual explanation as a plan. Your multi-agent epistemic planning work already gives a classical planner nested belief without an exponential blowup. Putting them together makes the explanation itself the output of an epistemic planner, where the goal is a belief state in the listener and the actions are things you can say. That is a different object from picking a minimal feature subset.

    The gap

    The explanation-planning probe returns model-reconciliation explanations (Chakraborti and colleagues; Vasileiou and colleagues for probabilistic scenarios), contrastive plan explanations through model restrictions, mental-model policies for sequential explanation, and recent agentic frameworks where a language model mediates the explanation dialogue. The listener model in all of them is a flat model difference or a learned policy. None runs a nested-belief epistemic planner to construct the explanation.

    First step

    With Francesco Leofante, whose counterfactual-scenarios work already treats counterfactuals as objects in a planning problem, encode a small explanation task as an epistemic planning instance with nested belief about the listener and compare the resulting explanation against a model-reconciliation baseline on the same domain.

    Bridgescs.AIcs.LOcs.HC
See it live

Run this live, Vaishak, on any researcher in any field.

What you just read is a snapshot. Inside Valency you walk the citation graph yourself, re-run the network, and chase these directions or wherever your own questions lead, in real time or async through your agents and your co-scientist. Join the waitlist and we’ll send you an invite.