election-methods@mailman.electorama.com

Technical discussion of election methods

View all threads

Test elections

RL
Richard Lung
Mon, Nov 29, 2021 10:38 PM

Hello Kristofer and All,

Can't at present find reply I did make to your test elections.

From memory, the nominal winner was B followed by C. But the real
winner was A because A had almost a quota -- to half a vote -- that was
not a statistically significant shortfall.

Your example reminded me of a methodological short-coming I had long
forgotten. I can't remember why I decided to ignore  it,  but it must
have been something along the lines that, for practical purposes, the
election count generally over-rides the exclusion count. In your example
it doesn't because it is a mere 3 candidate single- winner case -- the
case of minimal democracy. FAB STV is designed for a minimal 4 or 5 seat
case, and preferably more. I hadn't considered the hand count as a
possibly separate case.

I didn't even bother to pursue the option of statistical significance.
In your example, Formal winners B then C (technically) do not come any
where near the elective quota, at any statistical level of significance.
I may have been wrong not to discuss a significance requirement. It
depends on how actual binomial stv turns out.

A statistically significant binomial stv vote requires at least 30
voters ( approx 2^5 binomial  distribution). Thus your second example
cannot be given a significan count with binomial stv, which involves a
parametric statistic. Very small smples are only amenable to
non-parametric statistics.

However, I mentioned this issue of Kristofer very briefly at the end of
my latest publication:

https://www.smashwords.com/books/view/1116612

Am very tied-up for reasons already given.

Regards

Rchard Lung.

Hello Kristofer and All, Can't at present find reply I did make to your test elections. From memory, the nominal winner was B followed by C. But the real winner was A because A had almost a quota -- to half a vote -- that was not a statistically significant shortfall. Your example reminded me of a methodological short-coming I had long forgotten. I can't remember why I decided to ignore  it,  but it must have been something along the lines that, for practical purposes, the election count generally over-rides the exclusion count. In your example it doesn't because it is a mere 3 candidate single- winner case -- the case of minimal democracy. FAB STV is designed for a minimal 4 or 5 seat case, and preferably more. I hadn't considered the hand count as a possibly separate case. I didn't even bother to pursue the option of statistical significance. In your example, Formal winners B then C (technically) do not come any where near the elective quota, at any statistical level of significance. I may have been wrong not to discuss a significance requirement. It depends on how actual binomial stv turns out. A statistically significant binomial stv vote requires at least 30 voters ( approx 2^5 binomial  distribution). Thus your second example cannot be given a significan count with binomial stv, which involves a parametric statistic. Very small smples are only amenable to non-parametric statistics. However, I mentioned this issue of Kristofer very briefly at the end of my latest publication: https://www.smashwords.com/books/view/1116612 Am very tied-up for reasons already given. Regards Rchard Lung.
CC
Colin Champion
Thu, Dec 2, 2021 5:05 PM

To judge from the literature I would suppose that voting theory was part
of first-order logic or of graph theory, but it seems clear that the
right answer is Bayesian decision theory.

This is easiest to see under a spatial model, but I think it’s perfectly
general. We have vague prior information about the attributes of voters
and candidates (eg. their positions in space) and about voter behaviour
(how they will cast their ballots in the light of these attributes). We
can condition this prior information on the observations contained in a
set of ballots and thereby compute the posterior probability of any
desired function of the voters' and candidates' attributes. Our aim is
to identify the most representative candidate, where the degree of
representativeness can be expressed through a loss function, and where,
therefore, we will seek to identify the candidate whose posterior
expected loss is the least. The prior knowledge of voters' positions can
be thought of as a distribution of distributions, eg. the voters come
from a mixture of three identical and equally weighted circular
Gaussians whose means come from a further Gaussian.

This is a constructive approach which in principle might be used to
determine the winner of an election. We'd just need to integrate out all
the unknowns to find the expected losses of the candidates. This was in
essence the approach adopted by Good and Tideman in 1971 but not pursued
further. The same view of voting theory underlies the empirical
evaluations which have taken place subsequently: elections are sampled
under a vague prior, and the results are assessed under an appropriate
loss function. The only feature which conceals the decision-theoretic
basis is the persistent use of the term 'utility' where 'loss' would be
more correct.

Unfortunately the constructive approach seems to be numerically
intractable in cases of interest. If the number of voters was small, the
observations would provide probabilistic information which could be
integrated under the prior in the normal way. But as the number of
voters increases, the information becomes increasingly deterministic -
it degenerates to a set of equations. And therefore two cases arise.
Either the equations fully determine the parameters of the voter
distribution, in which case the prior almost drops out of the
calculation; or the equations constrain the voter parameters to a curved
manifold in which only the prior remains to be integrated.

The former case was encountered by Good and Tideman, which is why their
paper ended up as Bayesianism without the prior. Unfortunately their
model (a single Gaussian) is too simple to be of interest, given the
optimality of Condorcet methods under it.

I say that the prior 'almost' drops out of the calculation because Good
and Tideman's parameters have a degree of freedom which is independent
of the information in the ballots. This lies in the radial distance of
the three candidates from the centre of the circle whose circumference
they lie on. In general we may have prior information about this
distance, and it may affect the candidates' losses, so it seems a
suitable case for Bayesian treatment. But Good and Tideman adopt a
squared-distance loss function, and under this loss function (and this
function alone, I suspect) the radial distance is immaterial to the
identity of the optimal candidate. (The authors claim that the same
result applies to any loss function which depends solely on distance,
but I believe this to be an error.) Thus Good - of all people - made the
prior drop out completely.

It would be more interesting to adopt a Gaussian mixture prior as
sketched above. Even then we could perform nothing more than a
computational thought experiment. A truly realistic model would have to
allow for an arbitrary number of Gaussians in a space of any dimension,
and incorporate valence and random effects, and would need to allow for
tactical voting. It's hard to imagine any useful solution being obtainable.

But if voting theory is an insoluble problem in Bayesian decision
theory, then any voting method we encounter must be essentially ad hoc,
even if it draws on bomb-proof reasoning from another branch of
mathematics. At least we have the comfort of knowing that there is a
rigorous method of evaluating the solutions which are proposed to us.

CJC

To judge from the literature I would suppose that voting theory was part of first-order logic or of graph theory, but it seems clear that the right answer is Bayesian decision theory. This is easiest to see under a spatial model, but I think it’s perfectly general. We have vague prior information about the attributes of voters and candidates (eg. their positions in space) and about voter behaviour (how they will cast their ballots in the light of these attributes). We can condition this prior information on the observations contained in a set of ballots and thereby compute the posterior probability of any desired function of the voters' and candidates' attributes. Our aim is to identify the most representative candidate, where the degree of representativeness can be expressed through a loss function, and where, therefore, we will seek to identify the candidate whose posterior expected loss is the least. The prior knowledge of voters' positions can be thought of as a distribution of distributions, eg. the voters come from a mixture of three identical and equally weighted circular Gaussians whose means come from a further Gaussian. This is a constructive approach which in principle might be used to determine the winner of an election. We'd just need to integrate out all the unknowns to find the expected losses of the candidates. This was in essence the approach adopted by Good and Tideman in 1971 but not pursued further. The same view of voting theory underlies the empirical evaluations which have taken place subsequently: elections are sampled under a vague prior, and the results are assessed under an appropriate loss function. The only feature which conceals the decision-theoretic basis is the persistent use of the term 'utility' where 'loss' would be more correct. Unfortunately the constructive approach seems to be numerically intractable in cases of interest. If the number of voters was small, the observations would provide probabilistic information which could be integrated under the prior in the normal way. But as the number of voters increases, the information becomes increasingly deterministic - it degenerates to a set of equations. And therefore two cases arise. Either the equations fully determine the parameters of the voter distribution, in which case the prior almost drops out of the calculation; or the equations constrain the voter parameters to a curved manifold in which only the prior remains to be integrated. The former case was encountered by Good and Tideman, which is why their paper ended up as Bayesianism without the prior. Unfortunately their model (a single Gaussian) is too simple to be of interest, given the optimality of Condorcet methods under it. I say that the prior 'almost' drops out of the calculation because Good and Tideman's parameters have a degree of freedom which is independent of the information in the ballots. This lies in the radial distance of the three candidates from the centre of the circle whose circumference they lie on. In general we may have prior information about this distance, and it may affect the candidates' losses, so it seems a suitable case for Bayesian treatment. But Good and Tideman adopt a squared-distance loss function, and under this loss function (and this function alone, I suspect) the radial distance is immaterial to the identity of the optimal candidate. (The authors claim that the same result applies to any loss function which depends solely on distance, but I believe this to be an error.) Thus Good - of all people - made the prior drop out completely. It would be more interesting to adopt a Gaussian mixture prior as sketched above. Even then we could perform nothing more than a computational thought experiment. A truly realistic model would have to allow for an arbitrary number of Gaussians in a space of any dimension, and incorporate valence and random effects, and would need to allow for tactical voting. It's hard to imagine any useful solution being obtainable. But if voting theory is an insoluble problem in Bayesian decision theory, then any voting method we encounter must be essentially ad hoc, even if it draws on bomb-proof reasoning from another branch of mathematics. At least we have the comfort of knowing that there is a rigorous method of evaluating the solutions which are proposed to us. CJC
MK
Matthew Killebrew
Thu, Dec 2, 2021 5:07 PM

lol

On Thu, Dec 2, 2021 at 12:06 PM Colin Champion
colin.champion@routemaster.app wrote:

To judge from the literature I would suppose that voting theory was part
of first-order logic or of graph theory, but it seems clear that the
right answer is Bayesian decision theory.

This is easiest to see under a spatial model, but I think it’s perfectly
general. We have vague prior information about the attributes of voters
and candidates (eg. their positions in space) and about voter behaviour
(how they will cast their ballots in the light of these attributes). We
can condition this prior information on the observations contained in a
set of ballots and thereby compute the posterior probability of any
desired function of the voters' and candidates' attributes. Our aim is
to identify the most representative candidate, where the degree of
representativeness can be expressed through a loss function, and where,
therefore, we will seek to identify the candidate whose posterior
expected loss is the least. The prior knowledge of voters' positions can
be thought of as a distribution of distributions, eg. the voters come
from a mixture of three identical and equally weighted circular
Gaussians whose means come from a further Gaussian.

This is a constructive approach which in principle might be used to
determine the winner of an election. We'd just need to integrate out all
the unknowns to find the expected losses of the candidates. This was in
essence the approach adopted by Good and Tideman in 1971 but not pursued
further. The same view of voting theory underlies the empirical
evaluations which have taken place subsequently: elections are sampled
under a vague prior, and the results are assessed under an appropriate
loss function. The only feature which conceals the decision-theoretic
basis is the persistent use of the term 'utility' where 'loss' would be
more correct.

Unfortunately the constructive approach seems to be numerically
intractable in cases of interest. If the number of voters was small, the
observations would provide probabilistic information which could be
integrated under the prior in the normal way. But as the number of
voters increases, the information becomes increasingly deterministic -
it degenerates to a set of equations. And therefore two cases arise.
Either the equations fully determine the parameters of the voter
distribution, in which case the prior almost drops out of the
calculation; or the equations constrain the voter parameters to a curved
manifold in which only the prior remains to be integrated.

The former case was encountered by Good and Tideman, which is why their
paper ended up as Bayesianism without the prior. Unfortunately their
model (a single Gaussian) is too simple to be of interest, given the
optimality of Condorcet methods under it.

I say that the prior 'almost' drops out of the calculation because Good
and Tideman's parameters have a degree of freedom which is independent
of the information in the ballots. This lies in the radial distance of
the three candidates from the centre of the circle whose circumference
they lie on. In general we may have prior information about this
distance, and it may affect the candidates' losses, so it seems a
suitable case for Bayesian treatment. But Good and Tideman adopt a
squared-distance loss function, and under this loss function (and this
function alone, I suspect) the radial distance is immaterial to the
identity of the optimal candidate. (The authors claim that the same
result applies to any loss function which depends solely on distance,
but I believe this to be an error.) Thus Good - of all people - made the
prior drop out completely.

It would be more interesting to adopt a Gaussian mixture prior as
sketched above. Even then we could perform nothing more than a
computational thought experiment. A truly realistic model would have to
allow for an arbitrary number of Gaussians in a space of any dimension,
and incorporate valence and random effects, and would need to allow for
tactical voting. It's hard to imagine any useful solution being obtainable.

But if voting theory is an insoluble problem in Bayesian decision
theory, then any voting method we encounter must be essentially ad hoc,
even if it draws on bomb-proof reasoning from another branch of
mathematics. At least we have the comfort of knowing that there is a
rigorous method of evaluating the solutions which are proposed to us.

CJC

Election-Methods mailing list - see https://electorama.com/em for list
info

lol On Thu, Dec 2, 2021 at 12:06 PM Colin Champion <colin.champion@routemaster.app> wrote: > To judge from the literature I would suppose that voting theory was part > of first-order logic or of graph theory, but it seems clear that the > right answer is Bayesian decision theory. > > This is easiest to see under a spatial model, but I think it’s perfectly > general. We have vague prior information about the attributes of voters > and candidates (eg. their positions in space) and about voter behaviour > (how they will cast their ballots in the light of these attributes). We > can condition this prior information on the observations contained in a > set of ballots and thereby compute the posterior probability of any > desired function of the voters' and candidates' attributes. Our aim is > to identify the most representative candidate, where the degree of > representativeness can be expressed through a loss function, and where, > therefore, we will seek to identify the candidate whose posterior > expected loss is the least. The prior knowledge of voters' positions can > be thought of as a distribution of distributions, eg. the voters come > from a mixture of three identical and equally weighted circular > Gaussians whose means come from a further Gaussian. > > This is a constructive approach which in principle might be used to > determine the winner of an election. We'd just need to integrate out all > the unknowns to find the expected losses of the candidates. This was in > essence the approach adopted by Good and Tideman in 1971 but not pursued > further. The same view of voting theory underlies the empirical > evaluations which have taken place subsequently: elections are sampled > under a vague prior, and the results are assessed under an appropriate > loss function. The only feature which conceals the decision-theoretic > basis is the persistent use of the term 'utility' where 'loss' would be > more correct. > > Unfortunately the constructive approach seems to be numerically > intractable in cases of interest. If the number of voters was small, the > observations would provide probabilistic information which could be > integrated under the prior in the normal way. But as the number of > voters increases, the information becomes increasingly deterministic - > it degenerates to a set of equations. And therefore two cases arise. > Either the equations fully determine the parameters of the voter > distribution, in which case the prior almost drops out of the > calculation; or the equations constrain the voter parameters to a curved > manifold in which only the prior remains to be integrated. > > The former case was encountered by Good and Tideman, which is why their > paper ended up as Bayesianism without the prior. Unfortunately their > model (a single Gaussian) is too simple to be of interest, given the > optimality of Condorcet methods under it. > > I say that the prior 'almost' drops out of the calculation because Good > and Tideman's parameters have a degree of freedom which is independent > of the information in the ballots. This lies in the radial distance of > the three candidates from the centre of the circle whose circumference > they lie on. In general we may have prior information about this > distance, and it may affect the candidates' losses, so it seems a > suitable case for Bayesian treatment. But Good and Tideman adopt a > squared-distance loss function, and under this loss function (and this > function alone, I suspect) the radial distance is immaterial to the > identity of the optimal candidate. (The authors claim that the same > result applies to any loss function which depends solely on distance, > but I believe this to be an error.) Thus Good - of all people - made the > prior drop out completely. > > It would be more interesting to adopt a Gaussian mixture prior as > sketched above. Even then we could perform nothing more than a > computational thought experiment. A truly realistic model would have to > allow for an arbitrary number of Gaussians in a space of any dimension, > and incorporate valence and random effects, and would need to allow for > tactical voting. It's hard to imagine any useful solution being obtainable. > > But if voting theory is an insoluble problem in Bayesian decision > theory, then any voting method we encounter must be essentially ad hoc, > even if it draws on bomb-proof reasoning from another branch of > mathematics. At least we have the comfort of knowing that there is a > rigorous method of evaluating the solutions which are proposed to us. > > CJC > ---- > Election-Methods mailing list - see https://electorama.com/em for list > info >
RL
Richard Lung
Thu, Dec 2, 2021 7:13 PM

A voting distribution is a statistic. Representation is served by averages:
FAB STV: Four Averages Binomial Single Transferable Vote.

https://www.smashwords.com/books/view/806030

Profile page:
https://www.smashwords.com/profile/view/democracyscience

Richard Lung.

On 2 Dec 2021, at 5:05 pm, Colin Champion colin.champion@routemaster.app wrote:

To judge from the literature I would suppose that voting theory was part of first-order logic or of graph theory, but it seems clear that the right answer is Bayesian decision theory.

This is easiest to see under a spatial model, but I think it’s perfectly general. We have vague prior information about the attributes of voters and candidates (eg. their positions in space) and about voter behaviour (how they will cast their ballots in the light of these attributes). We can condition this prior information on the observations contained in a set of ballots and thereby compute the posterior probability of any desired function of the voters' and candidates' attributes. Our aim is to identify the most representative candidate, where the degree of representativeness can be expressed through a loss function, and where, therefore, we will seek to identify the candidate whose posterior expected loss is the least. The prior knowledge of voters' positions can be thought of as a distribution of distributions, eg. the voters come from a mixture of three identical and equally weighted circular Gaussians whose means come from a further Gaussian.

This is a constructive approach which in principle might be used to determine the winner of an election. We'd just need to integrate out all the unknowns to find the expected losses of the candidates. This was in essence the approach adopted by Good and Tideman in 1971 but not pursued further. The same view of voting theory underlies the empirical evaluations which have taken place subsequently: elections are sampled under a vague prior, and the results are assessed under an appropriate loss function. The only feature which conceals the decision-theoretic basis is the persistent use of the term 'utility' where 'loss' would be more correct.

Unfortunately the constructive approach seems to be numerically intractable in cases of interest. If the number of voters was small, the observations would provide probabilistic information which could be integrated under the prior in the normal way. But as the number of voters increases, the information becomes increasingly deterministic - it degenerates to a set of equations. And therefore two cases arise. Either the equations fully determine To judge from the literature I would suppose that voting theory was part of first-order logic or of graph theory, but it seems clear that the right answer is Bayesian decision theory.

This is easiest to see under a spatial model, but I think it’s perfectly general. We have vague prior information about the attributes of voters and candidates (eg. their positions in space) and about voter behaviour (how they will cast their ballots in the light of these attributes). We can condition this prior information on the observations contained in a set of ballots and thereby compute the posterior probability of any desired function of the voters' and candidates' attributes. Our aim is to identify the most representative candidate, where the degree of representativeness can be expressed through a loss function, and where, therefore, we will seek to identify the candidate whose posterior expected loss is the least. The prior knowledge of voters' positions can be thought of as a distribution of distributions, eg. the voters come from a mixture of three identical and equally weighted circular Gaussians whose means come from a further Gaussian.

This is a constructive approach which in principle might be used to determine the winner of an election. We'd just need to integrate out all the unknowns to find the expected losses of the candidates. This was in essence the approach adopted by Good and Tideman in 1971 but not pursued further. The same view of voting theory underlies the empirical evaluations which have taken place subsequently: elections are sampled under a vague prior, and the results are assessed under an appropriate loss function. The only feature which conceals the decision-theoretic basis is the persistent use of the term 'utility' where 'loss' would be more correct.

Unfortunately the constructive approach seems to be numerically intractable in cases of interest. If the number of voters was small, the observations would provide probabilistic information which could be integrated under the prior in the normal way. But as the number of voters increases, the information becomes increasingly deterministic - it degenerates to a set of equations. And therefore two cases arise. Either the equations fully determine the parameters of the voter distribution, in which case the prior almost drops out of the calculation; or the equations constrain the voter parameters to a curved manifold in which only the prior remains to be integrated.

The former case was encountered by Good and Tideman, which is why their paper ended up as Bayesianism without the prior. Unfortunately their model (a single Gaussian) is too simple to be of interest, given the optimality of Condorcet methods under it.

I say that the prior 'almost' drops out of the calculation because Good and Tideman's parameters have a degree of freedom which is independent of the information in the ballots. This lies in the radial distance of the three candidates from the centre of the circle whose circumference they lie on. In general we may have prior information about this distance, and it may affect the candidates' losses, so it seems a suitable case for Bayesian treatment. But Good and Tideman adopt a squared-distance loss function, and under this loss function (and this function alone, I suspect) the radial distance is immaterial to the identity of the optimal candidate. (The authors claim that the same result applies to any loss function which depends solely on distance, but I believe this to be an error.) Thus Good - of all people - made the prior drop out completely.

It would be more interesting to adopt a Gaussian mixture prior as sketched above. Even then we could perform nothing more than a computational thought experiment. A truly realistic model would have to allow for an arbitrary number of Gaussians in a space of any dimension, and incorporate valence and random effects, and would need to allow for tactical voting. It's hard to imagine any useful solution being obtainable.

But if voting theory is an insoluble problem in Bayesian decision theory, then any voting method we encounter must be essentially ad hoc, even if it draws on bomb-proof reasoning from another branch of mathematics. At least we have the comfort of knowing that there is a rigorous method of evaluating the solutions which are proposed to us.

CJC

Election-Methods mailing list - see https://electorama.com/em for list info

A voting distribution is a statistic. Representation is served by averages: FAB STV: Four Averages Binomial Single Transferable Vote. https://www.smashwords.com/books/view/806030 Profile page: https://www.smashwords.com/profile/view/democracyscience Richard Lung. On 2 Dec 2021, at 5:05 pm, Colin Champion <colin.champion@routemaster.app> wrote: To judge from the literature I would suppose that voting theory was part of first-order logic or of graph theory, but it seems clear that the right answer is Bayesian decision theory. This is easiest to see under a spatial model, but I think it’s perfectly general. We have vague prior information about the attributes of voters and candidates (eg. their positions in space) and about voter behaviour (how they will cast their ballots in the light of these attributes). We can condition this prior information on the observations contained in a set of ballots and thereby compute the posterior probability of any desired function of the voters' and candidates' attributes. Our aim is to identify the most representative candidate, where the degree of representativeness can be expressed through a loss function, and where, therefore, we will seek to identify the candidate whose posterior expected loss is the least. The prior knowledge of voters' positions can be thought of as a distribution of distributions, eg. the voters come from a mixture of three identical and equally weighted circular Gaussians whose means come from a further Gaussian. This is a constructive approach which in principle might be used to determine the winner of an election. We'd just need to integrate out all the unknowns to find the expected losses of the candidates. This was in essence the approach adopted by Good and Tideman in 1971 but not pursued further. The same view of voting theory underlies the empirical evaluations which have taken place subsequently: elections are sampled under a vague prior, and the results are assessed under an appropriate loss function. The only feature which conceals the decision-theoretic basis is the persistent use of the term 'utility' where 'loss' would be more correct. Unfortunately the constructive approach seems to be numerically intractable in cases of interest. If the number of voters was small, the observations would provide probabilistic information which could be integrated under the prior in the normal way. But as the number of voters increases, the information becomes increasingly deterministic - it degenerates to a set of equations. And therefore two cases arise. Either the equations fully determine To judge from the literature I would suppose that voting theory was part of first-order logic or of graph theory, but it seems clear that the right answer is Bayesian decision theory. This is easiest to see under a spatial model, but I think it’s perfectly general. We have vague prior information about the attributes of voters and candidates (eg. their positions in space) and about voter behaviour (how they will cast their ballots in the light of these attributes). We can condition this prior information on the observations contained in a set of ballots and thereby compute the posterior probability of any desired function of the voters' and candidates' attributes. Our aim is to identify the most representative candidate, where the degree of representativeness can be expressed through a loss function, and where, therefore, we will seek to identify the candidate whose posterior expected loss is the least. The prior knowledge of voters' positions can be thought of as a distribution of distributions, eg. the voters come from a mixture of three identical and equally weighted circular Gaussians whose means come from a further Gaussian. This is a constructive approach which in principle might be used to determine the winner of an election. We'd just need to integrate out all the unknowns to find the expected losses of the candidates. This was in essence the approach adopted by Good and Tideman in 1971 but not pursued further. The same view of voting theory underlies the empirical evaluations which have taken place subsequently: elections are sampled under a vague prior, and the results are assessed under an appropriate loss function. The only feature which conceals the decision-theoretic basis is the persistent use of the term 'utility' where 'loss' would be more correct. Unfortunately the constructive approach seems to be numerically intractable in cases of interest. If the number of voters was small, the observations would provide probabilistic information which could be integrated under the prior in the normal way. But as the number of voters increases, the information becomes increasingly deterministic - it degenerates to a set of equations. And therefore two cases arise. Either the equations fully determine the parameters of the voter distribution, in which case the prior almost drops out of the calculation; or the equations constrain the voter parameters to a curved manifold in which only the prior remains to be integrated. The former case was encountered by Good and Tideman, which is why their paper ended up as Bayesianism without the prior. Unfortunately their model (a single Gaussian) is too simple to be of interest, given the optimality of Condorcet methods under it. I say that the prior 'almost' drops out of the calculation because Good and Tideman's parameters have a degree of freedom which is independent of the information in the ballots. This lies in the radial distance of the three candidates from the centre of the circle whose circumference they lie on. In general we may have prior information about this distance, and it may affect the candidates' losses, so it seems a suitable case for Bayesian treatment. But Good and Tideman adopt a squared-distance loss function, and under this loss function (and this function alone, I suspect) the radial distance is immaterial to the identity of the optimal candidate. (The authors claim that the same result applies to any loss function which depends solely on distance, but I believe this to be an error.) Thus Good - of all people - made the prior drop out completely. It would be more interesting to adopt a Gaussian mixture prior as sketched above. Even then we could perform nothing more than a computational thought experiment. A truly realistic model would have to allow for an arbitrary number of Gaussians in a space of any dimension, and incorporate valence and random effects, and would need to allow for tactical voting. It's hard to imagine any useful solution being obtainable. But if voting theory is an insoluble problem in Bayesian decision theory, then any voting method we encounter must be essentially ad hoc, even if it draws on bomb-proof reasoning from another branch of mathematics. At least we have the comfort of knowing that there is a rigorous method of evaluating the solutions which are proposed to us. CJC ---- Election-Methods mailing list - see https://electorama.com/em for list info
KM
Kristofer Munsterhjelm
Thu, Dec 2, 2021 11:48 PM

On 12/2/21 6:05 PM, Colin Champion wrote:

To judge from the literature I would suppose that voting theory was part
of first-order logic or of graph theory, but it seems clear that the
right answer is Bayesian decision theory.

I would agree with you when the methods are being judged on a how good
it is according to e.g. how often it's susceptible to strategy, or how
often it returns the ideal winner, or its VSE results.

But I don't think the properties approach is quite as much Bayes stats.
Investigating what kind of properties are compatible or not is not
really a statistical thing: either a method fails (say) monotonicity or
it passes it. Trying to invent methods with new combinations of passing
criteria is... I would say, closer to linear algebra than anything,
although the exact linalg approach for doing so (what I almost did to
show unmanipulable majority and Condorcet as incompatible, and what
Craig Carey did to arrive at IFPP) doesn't really scale.

One could also argue that it's all economics, because public and social
choice originate in economics (namely, a way to apply economics to
political questions). Things like the median voter theorem and more
generally Hotelling's law feel very "econ-ish".

Unfortunately the constructive approach seems to be numerically
intractable in cases of interest. If the number of voters was small, the
observations would provide probabilistic information which could be
integrated under the prior in the normal way. But as the number of
voters increases, the information becomes increasingly deterministic -
it degenerates to a set of equations. And therefore two cases arise.
Either the equations fully determine the parameters of the voter
distribution, in which case the prior almost drops out of the
calculation; or the equations constrain the voter parameters to a curved
manifold in which only the prior remains to be integrated.

Let's say we use a spatial model. Then the parameters of that spatial
model would still be free (e.g. how many candidates, how many
dimensions, what kind of distribution do the candidates and the voters
follow). Deciding the values of these parameters might be like setting a
prior, because it doesn't necessarily arise from the empirical
distribution no matter how many voters you have.

For instance, let's say you want to figure out the Condorcet efficiency
of IRV. You use a spatial model based on current US politics, which has
a very high probability of being two-party. So you find out that
Plurality has near 100% Condorcet efficiency. But this is kind of
misleading, because the political situation is the way it is in order to
not make Plurality err to begin with.

I guess what I'm saying is that it doesn't seem like the hard parameters
scale badly with the number of voters; the difficult bit is instead how
to make sure the numbers are useful. The number of preferences does
indeed scale pretty badly, but I'd imagine Monte Carlo helps with this.

In some cases it might be possible to integrate exactly. One idea I've
had but never got around to implement is to make exact Yee maps. For a
particular point being the center of voter opinion (distributed
according to a Gaussian with some given variance), we can count the
proportion of say, A>B>C votes by taking the intersection of the polygon
of the Voronoi map designated closest to A, and the polygon of the
second-closest Voronoi map assigned to B, and the polygon of the
max-distance Voronoi map assigned to C. Then we can decompose the
polygon into triangles and integrate over them. This approach scales
badly in the number of candidates, but the number of voters are no
longer part of the picture; the fractions should be in the limit of
number of voters going to infinity (in practice some large finite number
due to numerical precision limits). Is that what you mean by the
equations fully determining the voter distribution?

-km

On 12/2/21 6:05 PM, Colin Champion wrote: > To judge from the literature I would suppose that voting theory was part > of first-order logic or of graph theory, but it seems clear that the > right answer is Bayesian decision theory. I would agree with you when the methods are being judged on a how good it is according to e.g. how often it's susceptible to strategy, or how often it returns the ideal winner, or its VSE results. But I don't think the properties approach is quite as much Bayes stats. Investigating what kind of properties are compatible or not is not really a statistical thing: either a method fails (say) monotonicity or it passes it. Trying to invent methods with new combinations of passing criteria is... I would say, closer to linear algebra than anything, although the exact linalg approach for doing so (what I almost did to show unmanipulable majority and Condorcet as incompatible, and what Craig Carey did to arrive at IFPP) doesn't really scale. One could also argue that it's all economics, because public and social choice originate in economics (namely, a way to apply economics to political questions). Things like the median voter theorem and more generally Hotelling's law feel very "econ-ish". > Unfortunately the constructive approach seems to be numerically > intractable in cases of interest. If the number of voters was small, the > observations would provide probabilistic information which could be > integrated under the prior in the normal way. But as the number of > voters increases, the information becomes increasingly deterministic - > it degenerates to a set of equations. And therefore two cases arise. > Either the equations fully determine the parameters of the voter > distribution, in which case the prior almost drops out of the > calculation; or the equations constrain the voter parameters to a curved > manifold in which only the prior remains to be integrated. Let's say we use a spatial model. Then the parameters of that spatial model would still be free (e.g. how many candidates, how many dimensions, what kind of distribution do the candidates and the voters follow). Deciding the values of these parameters might be like setting a prior, because it doesn't necessarily arise from the empirical distribution no matter how many voters you have. For instance, let's say you want to figure out the Condorcet efficiency of IRV. You use a spatial model based on current US politics, which has a very high probability of being two-party. So you find out that Plurality has near 100% Condorcet efficiency. But this is kind of misleading, because the political situation is the way it is in order to not make Plurality err to begin with. I guess what I'm saying is that it doesn't seem like the hard parameters scale badly with the number of voters; the difficult bit is instead how to make sure the numbers are useful. The number of preferences does indeed scale pretty badly, but I'd imagine Monte Carlo helps with this. In some cases it might be possible to integrate exactly. One idea I've had but never got around to implement is to make exact Yee maps. For a particular point being the center of voter opinion (distributed according to a Gaussian with some given variance), we can count the proportion of say, A>B>C votes by taking the intersection of the polygon of the Voronoi map designated closest to A, and the polygon of the second-closest Voronoi map assigned to B, and the polygon of the max-distance Voronoi map assigned to C. Then we can decompose the polygon into triangles and integrate over them. This approach scales badly in the number of candidates, but the number of voters are no longer part of the picture; the fractions should be in the limit of number of voters going to infinity (in practice some large finite number due to numerical precision limits). Is that what you mean by the equations fully determining the voter distribution? -km
KM
Kristofer Munsterhjelm
Fri, Dec 3, 2021 12:11 AM

On 11/29/21 11:38 PM, Richard Lung wrote:

Hello Kristofer and All,

Can't at present find reply I did make to your test elections.

From memory, the nominal winner was B followed by C. But the real
winner was A because A had almost a quota -- to half a vote -- that was
not a statistically significant shortfall.

Thank you for running one of the test elections through FAB. For the
benefit of the list, I'll repeat the election here :-)

752: A>B>C
1750: A>C>B
1752: B>C>A
1: C>A>B
750: C>B>A

With this test election, the outcome is A>B>C for every positional
system with alpha < 0.49 (where Plurality has alpha=0 and Antiplurality
is alpha=1), B>A>C for Borda, and C>B>A for every positional system with
alpha > 0.501.

I was trying to find out if single-winner FAB could be modeled as a
positional system elimination method (the way IRV is
Plurality-elimination). If it is, it's Coombs, because if the winner is
B, and C is second place, that means A must've been eliminated first.

The test is rather crude, though, so if you have more time later I can
try more refined tests. The second test election may also provide some
information.

I agree that three-candidate single-winner elections are not very good
demonstrators of a voting method that's intended to be multiwinner - but
the elections' simplicity also makes it easier to try experiments like
these.

-km

On 11/29/21 11:38 PM, Richard Lung wrote: > > > Hello Kristofer and All, > > > Can't at present find reply I did make to your test elections. > > From memory, the nominal winner was B followed by C. But the real > winner was A because A had almost a quota -- to half a vote -- that was > not a statistically significant shortfall. Thank you for running one of the test elections through FAB. For the benefit of the list, I'll repeat the election here :-) 752: A>B>C 1750: A>C>B 1752: B>C>A 1: C>A>B 750: C>B>A With this test election, the outcome is A>B>C for every positional system with alpha < 0.49 (where Plurality has alpha=0 and Antiplurality is alpha=1), B>A>C for Borda, and C>B>A for every positional system with alpha > 0.501. I was trying to find out if single-winner FAB could be modeled as a positional system elimination method (the way IRV is Plurality-elimination). If it is, it's Coombs, because if the winner is B, and C is second place, that means A must've been eliminated first. The test is rather crude, though, so if you have more time later I can try more refined tests. The second test election may also provide some information. I agree that three-candidate single-winner elections are not very good demonstrators of a voting method that's intended to be multiwinner - but the elections' simplicity also makes it easier to try experiments like these. -km
FS
Forest Simmons
Fri, Dec 3, 2021 4:44 AM

It's like the parable of the blind men and the elephant ... you see it from
the point(s) of view that are most familiar to you.

El jue., 2 de dic. de 2021 3:48 p. m., Kristofer Munsterhjelm <
km_elmet@t-online.de> escribió:

On 12/2/21 6:05 PM, Colin Champion wrote:

To judge from the literature I would suppose that voting theory was part
of first-order logic or of graph theory, but it seems clear that the
right answer is Bayesian decision theory.

I would agree with you when the methods are being judged on a how good
it is according to e.g. how often it's susceptible to strategy, or how
often it returns the ideal winner, or its VSE results.

But I don't think the properties approach is quite as much Bayes stats.
Investigating what kind of properties are compatible or not is not
really a statistical thing: either a method fails (say) monotonicity or
it passes it. Trying to invent methods with new combinations of passing
criteria is... I would say, closer to linear algebra than anything,
although the exact linalg approach for doing so (what I almost did to
show unmanipulable majority and Condorcet as incompatible, and what
Craig Carey did to arrive at IFPP) doesn't really scale.

One could also argue that it's all economics, because public and social
choice originate in economics (namely, a way to apply economics to
political questions). Things like the median voter theorem and more
generally Hotelling's law feel very "econ-ish".

Unfortunately the constructive approach seems to be numerically
intractable in cases of interest. If the number of voters was small, the
observations would provide probabilistic information which could be
integrated under the prior in the normal way. But as the number of
voters increases, the information becomes increasingly deterministic -
it degenerates to a set of equations. And therefore two cases arise.
Either the equations fully determine the parameters of the voter
distribution, in which case the prior almost drops out of the
calculation; or the equations constrain the voter parameters to a curved
manifold in which only the prior remains to be integrated.

Let's say we use a spatial model. Then the parameters of that spatial
model would still be free (e.g. how many candidates, how many
dimensions, what kind of distribution do the candidates and the voters
follow). Deciding the values of these parameters might be like setting a
prior, because it doesn't necessarily arise from the empirical
distribution no matter how many voters you have.

For instance, let's say you want to figure out the Condorcet efficiency
of IRV. You use a spatial model based on current US politics, which has
a very high probability of being two-party. So you find out that
Plurality has near 100% Condorcet efficiency. But this is kind of
misleading, because the political situation is the way it is in order to
not make Plurality err to begin with.

I guess what I'm saying is that it doesn't seem like the hard parameters
scale badly with the number of voters; the difficult bit is instead how
to make sure the numbers are useful. The number of preferences does
indeed scale pretty badly, but I'd imagine Monte Carlo helps with this.

In some cases it might be possible to integrate exactly. One idea I've
had but never got around to implement is to make exact Yee maps. For a
particular point being the center of voter opinion (distributed
according to a Gaussian with some given variance), we can count the
proportion of say, A>B>C votes by taking the intersection of the polygon
of the Voronoi map designated closest to A, and the polygon of the
second-closest Voronoi map assigned to B, and the polygon of the
max-distance Voronoi map assigned to C. Then we can decompose the
polygon into triangles and integrate over them. This approach scales
badly in the number of candidates, but the number of voters are no
longer part of the picture; the fractions should be in the limit of
number of voters going to infinity (in practice some large finite number
due to numerical precision limits). Is that what you mean by the
equations fully determining the voter distribution?

-km

Election-Methods mailing list - see https://electorama.com/em for list
info

It's like the parable of the blind men and the elephant ... you see it from the point(s) of view that are most familiar to you. El jue., 2 de dic. de 2021 3:48 p. m., Kristofer Munsterhjelm < km_elmet@t-online.de> escribió: > On 12/2/21 6:05 PM, Colin Champion wrote: > > To judge from the literature I would suppose that voting theory was part > > of first-order logic or of graph theory, but it seems clear that the > > right answer is Bayesian decision theory. > > I would agree with you when the methods are being judged on a how good > it is according to e.g. how often it's susceptible to strategy, or how > often it returns the ideal winner, or its VSE results. > > But I don't think the properties approach is quite as much Bayes stats. > Investigating what kind of properties are compatible or not is not > really a statistical thing: either a method fails (say) monotonicity or > it passes it. Trying to invent methods with new combinations of passing > criteria is... I would say, closer to linear algebra than anything, > although the exact linalg approach for doing so (what I almost did to > show unmanipulable majority and Condorcet as incompatible, and what > Craig Carey did to arrive at IFPP) doesn't really scale. > > One could also argue that it's all economics, because public and social > choice originate in economics (namely, a way to apply economics to > political questions). Things like the median voter theorem and more > generally Hotelling's law feel very "econ-ish". > > > Unfortunately the constructive approach seems to be numerically > > intractable in cases of interest. If the number of voters was small, the > > observations would provide probabilistic information which could be > > integrated under the prior in the normal way. But as the number of > > voters increases, the information becomes increasingly deterministic - > > it degenerates to a set of equations. And therefore two cases arise. > > Either the equations fully determine the parameters of the voter > > distribution, in which case the prior almost drops out of the > > calculation; or the equations constrain the voter parameters to a curved > > manifold in which only the prior remains to be integrated. > > Let's say we use a spatial model. Then the parameters of that spatial > model would still be free (e.g. how many candidates, how many > dimensions, what kind of distribution do the candidates and the voters > follow). Deciding the values of these parameters might be like setting a > prior, because it doesn't necessarily arise from the empirical > distribution no matter how many voters you have. > > For instance, let's say you want to figure out the Condorcet efficiency > of IRV. You use a spatial model based on current US politics, which has > a very high probability of being two-party. So you find out that > Plurality has near 100% Condorcet efficiency. But this is kind of > misleading, because the political situation is the way it is in order to > not make Plurality err to begin with. > > I guess what I'm saying is that it doesn't seem like the hard parameters > scale badly with the number of voters; the difficult bit is instead how > to make sure the numbers are useful. The number of preferences does > indeed scale pretty badly, but I'd imagine Monte Carlo helps with this. > > In some cases it might be possible to integrate exactly. One idea I've > had but never got around to implement is to make exact Yee maps. For a > particular point being the center of voter opinion (distributed > according to a Gaussian with some given variance), we can count the > proportion of say, A>B>C votes by taking the intersection of the polygon > of the Voronoi map designated closest to A, and the polygon of the > second-closest Voronoi map assigned to B, and the polygon of the > max-distance Voronoi map assigned to C. Then we can decompose the > polygon into triangles and integrate over them. This approach scales > badly in the number of candidates, but the number of voters are no > longer part of the picture; the fractions should be in the limit of > number of voters going to infinity (in practice some large finite number > due to numerical precision limits). Is that what you mean by the > equations fully determining the voter distribution? > > -km > ---- > Election-Methods mailing list - see https://electorama.com/em for list > info >
CC
Colin Champion
Fri, Dec 3, 2021 9:46 AM

I should have mentioned the Borda/Condorcet/Young school, which used a
jury model instead of a spatial model. Its central result is the
optimality of certain voting methods, including (in a limiting case) the
Borda count. The constructive methodology is maximum likelihood
estimation. Is this any more than a poor man's Bayesian decision theory?

CJC

On 03/12/2021 04:44, Forest Simmons wrote:

It's like the parable of the blind men and the elephant ... you see it
from the point(s) of view that are most familiar to you.

El jue., 2 de dic. de 2021 3:48 p. m., Kristofer Munsterhjelm
<km_elmet@t-online.de mailto:km_elmet@t-online.de> escribió:

I should have mentioned the Borda/Condorcet/Young school, which used a jury model instead of a spatial model. Its central result is the optimality of certain voting methods, including (in a limiting case) the Borda count. The constructive methodology is maximum likelihood estimation. Is this any more than a poor man's Bayesian decision theory? CJC On 03/12/2021 04:44, Forest Simmons wrote: > It's like the parable of the blind men and the elephant ... you see it > from the point(s) of view that are most familiar to you. > > El jue., 2 de dic. de 2021 3:48 p. m., Kristofer Munsterhjelm > <km_elmet@t-online.de <mailto:km_elmet@t-online.de>> escribió: >