I started drafting a post on this subject, but it got longer and longer
so I turned it into a web page:
http://www.masterlyinactivity.com/condorcet/semantics.html
Briefly, I argue that discussions of voting methods are only
meaningful if a semantics can be attached to the correctness of
electoral decisions; that models (such as jury and spatial models) can
provide such a semantics, leading to a Bayesian interpretation of
correctness; that the logical criteria stand or fall according to
whether they can be validated under a suitable semantics; that some
stand, some fall, and some are best seen as statistical approximations.
The topics I discuss are ones I have not seen addressed elsewhere. I
have no idea how new my ideas are, or whether, if I was better grounded
in the field, I'd have been able to discuss the subject with greater wisdom.
CJC
My web browser doesn't support it. I look forward to reading what you have
so far. The more we can identify what we are trying to do, the better ... a
truly laudable endeavor!
El mar., 1 de mar. de 2022 9:04 a. m., Colin Champion <
colin.champion@routemaster.app> escribió:
I started drafting a post on this subject, but it got longer and longer
so I turned it into a web page:
http://www.masterlyinactivity.com/condorcet/semantics.html
Briefly, I argue that discussions of voting methods are only
meaningful if a semantics can be attached to the correctness of
electoral decisions; that models (such as jury and spatial models) can
provide such a semantics, leading to a Bayesian interpretation of
correctness; that the logical criteria stand or fall according to
whether they can be validated under a suitable semantics; that some
stand, some fall, and some are best seen as statistical approximations.
The topics I discuss are ones I have not seen addressed elsewhere. I
have no idea how new my ideas are, or whether, if I was better grounded
in the field, I'd have been able to discuss the subject with greater
wisdom.
Election-Methods mailing list - see https://electorama.com/em for list
info
Hi Forest and Colin,
Forest, does your browser support this link?
https://www.masterlyinactivity.com/condorcet/semantics.html
Note the difference between "http" (bad old insecure link) and "https"
(secure, encrypted link). These days, Chrome and other browsers are
starting to require "https", even when you don't care about
security/privacy/etc.
Colin: I suppose that's useful for you to know too, since your website
supports "https" links as a replacement for "http" links.
Rob
Rob
On Tue, Mar 1, 2022 at 4:37 PM Forest Simmons
forest.simmons21@gmail.com wrote:
My web browser doesn't support it. I look forward to reading what you have so far. The more we can identify what we are trying to do, the better ... a truly laudable endeavor!
El mar., 1 de mar. de 2022 9:04 a. m., Colin Champion colin.champion@routemaster.app escribió:
I started drafting a post on this subject, but it got longer and longer
so I turned it into a web page:
http://www.masterlyinactivity.com/condorcet/semantics.html
Briefly, I argue that discussions of voting methods are only
meaningful if a semantics can be attached to the correctness of
electoral decisions; that models (such as jury and spatial models) can
provide such a semantics, leading to a Bayesian interpretation of
correctness; that the logical criteria stand or fall according to
whether they can be validated under a suitable semantics; that some
stand, some fall, and some are best seen as statistical approximations.
The topics I discuss are ones I have not seen addressed elsewhere. I
have no idea how new my ideas are, or whether, if I was better grounded
in the field, I'd have been able to discuss the subject with greater wisdom.
Election-Methods mailing list - see https://electorama.com/em for list info
Election-Methods mailing list - see https://electorama.com/em for list info
On 01.03.2022 18:03, Colin Champion wrote:
I started drafting a post on this subject, but it got longer and longer
so I turned it into a web page:
http://www.masterlyinactivity.com/condorcet/semantics.html
Briefly, I argue that discussions of voting methods are only
meaningful if a semantics can be attached to the correctness of
electoral decisions; that models (such as jury and spatial models) can
provide such a semantics, leading to a Bayesian interpretation of
correctness; that the logical criteria stand or fall according to
whether they can be validated under a suitable semantics; that some
stand, some fall, and some are best seen as statistical approximations.
The topics I discuss are ones I have not seen addressed elsewhere. I
have no idea how new my ideas are, or whether, if I was better grounded
in the field, I'd have been able to discuss the subject with greater
wisdom.
The closest thing I know of are the frequentist approaches to modeling
election methods, e.g. Kemeny as MLE of a certain extension of the model
in Condorcet's jury theorem, and generalizations to this approach (e.g.
https://www.jstor.org/stable/41106629).
Bayesian statistics is not my field, but as I understand your page,
you're trying to show whether certain models can naturally result in
monotonicity and participation failures, similar to how e.g. IIA by
necessity arises from every ordinal method.
The idea of trying to recover the statistical parameters of issue space
and then electing the best candidate is a good one; I think the reason
there hasn't been much of it is that voting methods research has been
more focused on strategy and on pass-or-fail criteria.
Do there exist extensions to Bayesian stats that handle cases where the
input data may be adversarially corrupted by some party who seeks to
confuse the process? That'd be like voting method strategy.
There's also the issue of noise. In your tetrahedron vs line example,
it's possible that one of the voters filled in the ballot incorrectly
(or misjudged or something). But I imagine that "ordinary" Bayesian
stats could handle noise with appropriately broad priors.
I also think that some properties are considered as "embarrassment
criteria", as in: it seems illogical that a method should fail this
particular criterion, and the opponents of the method might use this to
ridicule the method, so we better patch it up. Or it may be part of how
we think a voting method should behave, no matter what.
Binary properties give a certain guarantee that if something strange
happens, it won't be of this particular type. Let's say, for instance,
that we want to generalize majority rule. And suppose that under a
particular model (jury, say), Borda is optimal. Then if we want to
uphold majority rule no matter what, but still want some performance on
jury models, then Black may be a better choice than Borda.
In a way, there's always a question of what matters.
But looking closely into how VSE might be optimized by Bayesian methods
and how certain spaces imply certain properties is definitely
worthwhile, I think :-)
-km
Hi Colin,
Le mardi 1 mars 2022, 11:04:13 UTC−6, Colin Champion colin.champion@routemaster.app a écrit :
I started drafting a post on this subject, but it got longer and longer
so I turned it into a web page:
http://www.masterlyinactivity.com/condorcet/semantics.html
Briefly, I argue that discussions of voting methods are only
meaningful if a semantics can be attached to the correctness of
electoral decisions; that models (such as jury and spatial models) can
provide such a semantics, leading to a Bayesian interpretation of
correctness; that the logical criteria stand or fall according to
whether they can be validated under a suitable semantics; that some
stand, some fall, and some are best seen as statistical approximations.
The topics I discuss are ones I have not seen addressed elsewhere. I
have no idea how new my ideas are, or whether, if I was better grounded
in the field, I'd have been able to discuss the subject with greater wisdom.
I read through this a few times. I think I understand most of it, other than the
demonstrations, which will probably be my own failing.
It seems like your main stance is that even if we assume that all votes are
sincere, we should still have a theory about where the rankings come from, if we
want to talk about how a method ought to behave.
Maybe this line of thinking can be seen in Yee diagrams, where no strategy is
considered and methods are judged as to whether win regions match the Voronoi
diagram.
I may not understand free variables vs. bound variables. It sounds like with
bound variables, some standard is "right" if it tends to be right. While free
variables would judge standards to be right or wrong in every specific case.
So you say that a unanimity (or Pareto) criterion would be unjustified under
free variables with a jury/valence model. While I can see that most standards
might be unusable with free variables, in the case of unanimity I don't see how
it could be unjustified. When you discuss IIA it sounds like the jury/valence
model is essentially an underlying ranking for each voter, with no issue space.
Is there an assumption that a voter's ranking can be wrong in some sense?
If so, does that also apply to the spatial model? (i.e. that the voter is not
placed correctly in space.)
I think it's curious that when you bring in "external facts," in both ways
you've done this, those facts would tell you who the best candidate is in any
given case. There's no ambiguity. Could we have external facts that don't
necessarily do this? That wouldn't necessarily be useless, since we could at
least discuss phenomena in the new terms instead of just on the cast ballots.
What if we think that no model of external facts is justifiable in some
environment? Can we say nothing then? Maybe our conclusions would reveal some
implicit assumptions, I suppose, indicating something about what we suspect the
external facts are.
If the model must tell us the winner (due to one's definition of what a model
is), I wonder what stops a Condorcet advocate from copy/pasting the cast ballots
as external facts and declaring that the Condorcet winner within the external
facts is the targeted, right winner.
When you discuss the validity of Participation under a spatial model, do you
consider whether the transformations envisioned by the criterion can actually
be achieved under the model? Perhaps there could be a gray area of satisfaction
where we say a method fails a criterion, but only if the underlying model is
wrong (or e.g. if voters are insincere).
I don't understand your criticism of the notion that Participation "relates ...
to additional incentives offered to voters." Some of us, and Woodall, do
recognize Participation as a monotonicity criterion, for what it's worth. But I
don't know how to explain the strategic implications except in terms of
incentives offered to voters.
When you discuss IIA I am a little confused. You seem very against it, but only
discuss the jury/valence model. Is it simply obvious that it also applies to the
spatial model? What you say is that it's "clearly invalid under a Bayesian
semantics" which seems like an even broader claim.
In order to discuss IRV, you say that if we are willing to accept monotonicity
as a "valid" criterion under the model, then we can use it to judge the accuracy
of methods.
I expect there are three possibilities for a proposed criterion and model.
Either the criterion is valid (i.e. necessarily true), or it's incompatible, or
it's orthogonal, non-contradictory. In the case of Participation under a spatial
model, you say this "cannot be valid." I guess you showed it's incompatible.
What do you suppose is the status of a criterion in the middle state, neither
incompatible nor necessarily true? Must we be indifferent to it, having
exhausted the only legitimate method of assessing it?
Somewhat related to this, you point out that Borda is ideal or near ideal under
the jury model. But almost no one actually advocates Borda, for, let's say,
reasons of strategy. The model can't speak to these reasons, or at least, can't
ultimately offer us anything but an insistence that we use a Borda-like method.
I find that disturbing, since the model is what we're using to gauge potential
properties and that seems like its whole point.
Kristofer mentioned "embarrassment criteria" as an issue. I will often consider
a criterion important because I think that other people think it's important.
This results in a lot of subjectivity. For example maybe we can quantify
monotonicity failure, but how much is too much? Difficult to see a way out of
this.
Kevin
Kevin - you make a lot of points, I will try to reply to a few of them.
I'm sorry if I didn't make myself clear. My own view is that the
correctness of electoral decisions needs to be understood evidentially,
which in the end means Bayesianly. I am pretty sure that Jack Good held
this view and that Kenneth Arrow did not; but when I try to formulate
some alternative understanding of the correctness of electoral decisions
such as may have been in Arrow's head, I honestly struggle. I cannot be
entirely clear about what a non-evidential semantics would look like.
I would not say, as a philosophical position, that the semantics
need to make reference to a model of voting. If it was possible to put
together a model-free formalistic semantics, it might well be
acceptable. But in practice I cannot see any viable semantics which does
not have a model underlying it. Take this as a challenge, not as a dogma.
My notion of a jury model is that each candidate has an objective
valence (or excellence), which we can take to be gaussianly distributed,
and that each voter ranks candidates according to his or her own noisy
estimates of their valences. This is essentially the model Peyton Young
used, except that he worked with probabilities rather than with
statistical distributions. (I prefer my own approach because I find it
hard to keep a clear head when dealing with probabilities. I'm not sure
Young himself entirely succeeded - I think at one point he says
"independent" when he means "conditionally independent given... ".)
I'm quite struck by my counterexample to IIA (49% A>C>B, 51% B>A>C,
which is simply a numerical version of Good's argument). It seems to me
obvious that A is the rightful winner, and that if C is removed from the
ballots then B becomes the rightful winner. Certainly anyone who thinks
that C simply drops out of the analysis, and that my example is
equivalent to 49% A>B, 51% B>A is making an elementary statistical
error. I cannot believe that Arrow would have made such a mistake, so I
conclude that he understood electoral correctness in a different sense
than Good and I do. If only he had told us what it was!
In fact I haven't verified my numbers. Given a little time I could make
a plot, similar to the ones on my web page, of isofactors for the pair
of fractional ballots 0.49@A>C>B, 0.51@B>A>C. I assume there would be a
hill whose summit was in the region in which A is better than B and B is
better than C. The more ballots you accumulate, the steeper the hill
becomes.
Nor have I thought through the application of IIA to spatial models. I
don't see any reason why it should apply, and maybe Arrow's proof
makes it unnecessary to fill in the details. After all, Arrow never
limited his theorem to spatial models, and one counterexample is all
that's normally called for.
I agree that the Borda count is hopeless in the presence of tactical
voting, even under a jury model. But philosophically I don't feel
threatened by this. My view is that the right electoral decision is the
one which is most likely to give the best candidate, or whose expected
loss is least, or whatever, given - simultaneously - a model of sincere
voting behaviour and a model of how voters try to beat the system. In
practice I recognise that the calculation is beyond anything I can
envisage performing, and that even partial results are likely to
constitute a significant advance. Under a jury model with tactical
voting, I have some evidence that the best methods are the ones with the
worst reputations: FPTP and IRV and their extensions (including
Condorcet/Hare).
I think my words about the logical criteria may have come across as more
dogmatic than I intended. I hadn't realised that participation was
sometimes recognised as a form of monotonicity. I think I would say that
some criteria are always and necessarily true; some are always or nearly
always true, and may may provide useful guidance; and some are totally
untrustworthy. So long as one doesn't treat a statistical approximation
as a logical truth I have no real complaint.
Colin
I wrote in another thread that "under a jury model with tactical voting,
I have some evidence that the best methods are the ones with the worst
reputations: FPTP and IRV and their extensions". Perhaps some people
would like to see the evidence. The following table gives mean valence
losses (so smaller is better) under a jury model with voters attempting
a burial strategy. It has to be viewed in a fixed-width font.
random fptp sptp av sinkhorn borda mj coombs
115.6767 4.2340 15.4462 3.4934 22.3533 21.3425 - 36.3451
condorcet benham btr nanson minimax minisum rp river
schulze asm
- 3.5469 4.6945 9.5228 7.7089 8.0696 8.4014
8.3864 9.6397 9.8513
condorcet+ random fptp sptp av borda
9.9406 4.9425 11.5035 3.5441 14.2516
copeland+ random fptpf fptpr sptp avf avr bordaf
bordar minimaxfminimaxr
11.4423 9.6349 10.7388 12.4150 9.5974 10.5371 14.4200
12.7967 8.7613 12.0312
smith+ random fptpf fptpr sptp avf avr bordaf
bordar minimaxfminimaxrtideman q&dc
9.6397 4.9559 4.9856 11.5026 3.5442 3.5532 14.2516
10.0643 7.7106 7.7106 4.0854 7.3363
AV (=IRV) seems to be best, while most Condorcet methods do badly and
the Borda count does appallingly.
These results are from my own evaluation software; full details are at
https://www.masterlyinactivity.com/condorcet/condorcet.html The call is
"condorcet 5 101 100000 jury:2+bur". Obviously the correctness of my
code is not guaranteed.
I've never seen an evaluation under a jury model, even assuming sincere
voting. I certainly don't rate such models highly, but I think
discussions of voting need a model and spatial models aren't the whole
truth, so it helps to keep alternatives in mind.
CJC
On 06.03.2022 14:41, Colin Champion wrote:
I wrote in another thread that "under a jury model with tactical voting,
I have some evidence that the best methods are the ones with the worst
reputations: FPTP and IRV and their extensions". Perhaps some people
would like to see the evidence. The following table gives mean valence
losses (so smaller is better) under a jury model with voters attempting
a burial strategy. It has to be viewed in a fixed-width font.
That makes sense because FPTP and IRV have in common that burial doesn't
really work: either it does nothing or A>W voters burying W makes
someone else than A win. And it also seems reasonable that the minmax
methods do worse than Condorcet-IRV hybrids since the latter pass DMTBR
and the former don't.
For compromising strategy, I imagine that FPTP and IRV would be
considerably worse than the Condorcet methods.
I suppose you could try Daniel Carrera's trivial strategy, but I'm not
sure how you would translate it into a score. A method is susceptible to
trivial strategy in an election if there exists a candidate X different
from the winner W so that when all the voters who prefer X to W move X
first on their ballots and W last, then X wins.
But if the strategy can be pulled off for more than one X, then what
election counts for calculating the score of the method? I'm not sure.
I'd guess that the trivial strategy would also put the IRV-likes ahead
of (better than) the rest.
What's most surprising, IMHO, is that Condorcet-IRV hybrids do worse
than just plain IRV. James Green-Armytage gives a result where
Condorcetifying a method (electing the CW if there is one) can't
increase the proportion of elections where strategy works, if a majority
can force an outcome for the method in question. But apparently this
does not extend to gracefully degrading: your results seem to show that
even if there are fewer elections where strategy is possible, then under
this model, the strategy does more harm on average where it is possible
to execute.
random fptp sptp av sinkhorn borda mj coombs
115.6767 4.2340 15.4462 3.4934 22.3533 21.3425 - 36.3451
condorcet benham btr nanson minimax minisum rp river
schulze asm
- 3.5469 4.6945 9.5228 7.7089 8.0696 8.4014
8.3864 9.6397 9.8513
condorcet+ random fptp sptp av borda
9.9406 4.9425 11.5035 3.5441 14.2516
copeland+ random fptpf fptpr sptp avf avr bordaf bordar
minimaxfminimaxr
11.4423 9.6349 10.7388 12.4150 9.5974 10.5371 14.4200
12.7967 8.7613 12.0312
smith+ random fptpf fptpr sptp avf avr bordaf bordar
minimaxfminimaxrtideman q&dc
9.6397 4.9559 4.9856 11.5026 3.5442 3.5532 14.2516
10.0643 7.7106 7.7106 4.0854 7.3363
Some of the columns seem to almost run into each other; perhaps it would
be better to split them over two rows.
That Copeland does so badly also seems to tally well with my own
results. If I'm correct, then Landau should be slightly worse than
Smith, but not anywhere near Copeland.
Are the Smith+ methods Smith// or Smith,?
AV (=IRV) seems to be best, while most Condorcet methods do badly and
the Borda count does appallingly.
These results are from my own evaluation software; full details are at
https://www.masterlyinactivity.com/condorcet/condorcet.html The call is
"condorcet 5 101 100000 jury:2+bur". Obviously the correctness of my
code is not guaranteed.
I've never seen an evaluation under a jury model, even assuming sincere
voting. I certainly don't rate such models highly, but I think
discussions of voting need a model and spatial models aren't the whole
truth, so it helps to keep alternatives in mind.
Another non-spatial/jury-like possibility is the Kemeny model: there's a
distinguished "ground truth" order of the candidates. Each voter casts a
vote according to this ground truth, but for every pairwise preference,
flips them around with some probability p.
So the procedure would be to take the ground truth, and rejection sample
applying the independent noise to the pairwise components, until the
result is transitive (i.e. no voter can cast a cyclical ballot, so just
keep trying until it isn't).
E.g. if the G.T. is A>B>C>D, then a ballot may become
A>B A>C A>D B>C B>D C>D
truth T T T T T T
corrupted T F T F T T
C>A>B>D, instead.
Then the voting method's performance is given by the Kendall tau
distance between the ground truth and the social order produced by the
method when given these ballots.
And if you want to assign scores to impartial culture, then Durand's
hypersphere model might work. I'm not entirely sure how that works (or
how you would sample from it), but Forest showed that it's a natural
extension of impartial culture.
-km
Kristofer - for completeness here are my other results. Compromising, as
you imply, is much the same as sincere voting; false cycles are much
like burial. My final condition is "exhaustive" tactical voting, in
which all voters supporting a given candidate submit the same insincere
ballot. A subversion is considered successful if it leads to the
insincere voters getting their way and their candidate is worse than the
sincere winner; and the loss is the resulting drop in valence. I'm not
sure how close this is to Daniel's tests, but it was based on what he
was saying at the time. Exhaustive tactical voting is again much like
burial.
Other parameters are the same as in my previous run.
Like you, I was a little perturbed by Condorcet/Hare not being at
least as good as IRV; I simply assumed (or hoped) that the conditions of
JGA's theorem didn't hold.
Colin
Compromising
random fptp sptp av sinkhorn borda mj coombs
115.6767 4.2340 17.5522 3.4865 2.8701 2.8441 - 3.4740
condorcet benham btr nanson minimax minisum rp river
schulze asm
- 3.4419 3.3958 3.3815 3.3559 3.3530 3.3659
3.3858 3.5267 3.5209
condorcet+ random fptp sptp av borda
5.9255 3.3738 3.7981 3.4409 3.3058
copeland+ random fptpf fptpr sptp avf avr bordaf
bordar minimaxfminimaxr
3.4407 3.3690 3.3868 3.4452 3.4296 3.4302 3.3063
3.3388 3.3530 3.3647
smith+ random fptpf fptpr sptp avf avr bordaf
bordar minimaxfminimaxrtideman q&dc
3.5267 3.3720 3.3890 3.6909 3.4409 3.4382 3.3058
3.3259 3.3559 3.3559 3.4344 3.5345
False cycles
random fptp sptp av sinkhorn borda mj coombs
115.6767 4.2340 17.5522 3.4932 10.2714 10.0818 - 26.3221
condorcet benham btr nanson minimax minisum rp river
schulze asm
- 3.5519 5.4141 8.5122 8.6953 8.5354 8.7293
8.7462 12.4334 9.2892
condorcet+ random fptp sptp av borda
11.1509 5.2152 6.1789 3.5529 7.3317
copeland+ random fptpf fptpr sptp avf avr bordaf
bordar minimaxfminimaxr
10.6530 5.1834 5.3717 6.0292 3.7368 3.7725 7.1250
10.5974 8.3103 8.3840
smith+ random fptpf fptpr sptp avf avr bordaf
bordar minimaxfminimaxrtideman q&dc
12.4334 5.2357 5.4327 6.1769 3.5529 3.5543 7.3305
10.8659 8.7099 8.7099 4.1473 12.2860
Exhaustive
random fptp sptp av sinkhorn borda mj coombs
115.6767 4.2340 17.5522 3.4934 29.7429 28.3067 - 47.0813
condorcet benham btr nanson minimax minisum rp river
schulze asm
- 3.5522 5.4275 11.0898 10.0524 10.7085 11.5937
11.6862 15.6366 9.9130
condorcet+ random fptp sptp av borda
18.3528 5.2430 13.4912 3.5582 20.3795
copeland+ random fptpf fptpr sptp avf avr bordaf
bordar minimaxfminimaxr
15.7254 10.6274 10.7833 15.1104 10.3921 10.5389 19.4887
13.4124 12.2174 12.2592
smith+ random fptpf fptpr sptp avf avr bordaf
bordar minimaxfminimaxrtideman q&dc
15.6366 5.3535 6.3920 13.5263 3.5589 3.5666 20.3782
17.8307 10.0525 10.0525 4.1848 19.7032
Hi Colin,
My notion of a jury model is that each candidate has an objective
valence (or excellence), which we can take to be gaussianly distributed,
and that each voter ranks candidates according to his or her own noisy
estimates of their valences. This is essentially the model Peyton Young
used, except that he worked with probabilities rather than with
statistical distributions. (I prefer my own approach because I find it
hard to keep a clear head when dealing with probabilities. I'm not sure
Young himself entirely succeeded - I think at one point he says
"independent" when he means "conditionally independent given... ".)
Ok. A sort of hidden rating, and the voters use them to populate whichever type
of ballot is provided.
I'm quite struck by my counterexample to IIA (49% A>C>B, 51% B>A>C,
which is simply a numerical version of Good's argument). It seems to me
obvious that A is the rightful winner, and that if C is removed from the
ballots then B becomes the rightful winner. Certainly anyone who thinks
that C simply drops out of the analysis, and that my example is
equivalent to 49% A>B, 51% B>A is making an elementary statistical
error. I cannot believe that Arrow would have made such a mistake, so I
conclude that he understood electoral correctness in a different sense
than Good and I do. If only he had told us what it was!
Doesn't IIA hold within a jury model? I mean the estimates of valences.
Perhaps Arrow intends such a model, and means that it would be reasonable to
hope that a property of the model could also be reflected in the procedure.
I agree that the Borda count is hopeless in the presence of tactical
voting, even under a jury model. But philosophically I don't feel
threatened by this. My view is that the right electoral decision is the
one which is most likely to give the best candidate, or whose expected
loss is least, or whatever, given - simultaneously - a model of sincere
voting behaviour and a model of how voters try to beat the system.
My view is pretty similar.
Kevin