Being an electrical engineer that was ABD for a PhD in communications systems and signal processing, I have a little trouble seeing the connection to Shannon Information Theory. Either in the measure of information content of a message or set of messages or of the definition of entropy or of the capacity of a channel to carry information.So could someone make the connection for me? I get that set of ordinal ballot data is discrete information and there's some way, such as Huffman coding, to represent that information in the most compact and essential way possible.But I don't see the connection to social choice theory. Can someone help?robertPowered by Cricket Wireless------ Original message------From: Forest SimmonsDate: Sat, Jun 4, 2022 4:06 PMTo: Richard Lung;Cc: EM;Subject:Re: [EM] ThermodynamicsTrue!Do an internet search of "information mechanics" to confirm the validity of this tight connection.Information mechanics seems to be the key for the "unified field theory" Einstei
n was looking for ... and more ... unification of classical and quantum fields for all of the forces ... strong, weak, and intermediate... if not a "theory of everything."El sáb., 4 de jun. de 2022 6:26 a. m., Richard Lung voting@ukscientists.com escribió:
Forest,
The
efficiency of heat engines, in thermodynamics, offer an analogy
with voting
methods. Many other sciences do so, if voting method follows the
Stevens
structure of measurement, held in common by other branches of
science. (I
published a free e-book, about scientific models of election
method, called:
Science is Ethics as Electics.)
The
basic principle, that thermodynamics and election method have in
common is
conservation, either of energy or information. (I believe
scientists are
currently translating energy terms into information terms.)
Common-place
teachings of social choice theory, including the American
Mathematics Society,
usually make the claim that there is no perfect voting system.
The equivalent
statement in thermodynamics is that there is no perpetual motion
machine.
As
you point out, that does not preclude voting methods of
different efficiency,
the equivalent of heat engines of differing efficiency. The
engines depend on
efficient transfer of surplus heat, to work requirements, to
keep the engine
going. Similarly, transfers of vote surpluses, to elective
quotas, keep the
count procedure going. Heat forms a random distribution of
motion. And votes
typically form a random distribution of choice (subject to left
or right
skews).
Binomial STV
would perhaps
be rather more efficient than traditional STV, because it
rationally conserves
exclusion information. In rough analogy, a binomial STV “heat
engine” is better
“insulated,” to conserve heat. Thermodynamics is not just a
dynamic of heat but
also its insulation, in a closed system. Likewise, an election
method is not
just an active election, but also a closed system of exclusion.
Regards,
Richard
Lung.
As a software guy, the connection I make is to something I have wondered
for a while -- whether it is useful to study social choice functions as
lossy compression algorithms. I haven't thought it through, but it could be
interesting to see if the rate-distortion branch of information theory
would apply.
On Sat, Jun 4, 2022, 9:09 PM robert bristow-johnson <
rbj@audioimagination.com> wrote:
Being an electrical engineer that was ABD for a PhD in communications
systems and signal processing, I have a little trouble seeing the
connection to Shannon Information Theory. Either in the measure of
information content of a message or set of messages or of the definition of
entropy or of the capacity of a channel to carry information.
So could someone make the connection for me?
I get that set of ordinal ballot data is discrete information and there's
some way, such as Huffman coding, to represent that information in the most
compact and essential way possible.
But I don't see the connection to social choice theory. Can someone help?
robert
Powered by Cric ket Wireless
------ Original message------
*From: *Forest Simmons
*Date: *Sat, Jun 4, 2022 4:06 PM
*To: *Richard Lung;
*Cc: *EM;
*Subject:*Re: [EM] Thermodynamics
True!
Do an internet search of "information mechanics" to confirm the validity
of this tight connection.
Information mechanics seems to be the key for the "unified field theory"
Einstein was looking for ... and more ... unification of classical and
quantum fields for all of the forces ... strong, weak, and intermediate...
if not a "theory of everything."
El sáb., 4 de jun. de 2022 6:26 a. m., Richard Lung <
voting@ukscientists.com> escribió:
Forest,
The efficiency of heat engines, in thermodynamics, offer an analogy with
voting methods. Many other sciences do so, if voting method follows the
Stevens structure of measurement, held in common by other branches of
science. (I published a free e-book, about scientific models of election
method, called: Science is Ethics as Electics.)
The basic principle, that thermodynamics and election method have in
common is conservation, either of energy or information. (I believe
scientists are currently translating energy terms into information terms.)
Common-place teachings of social choice theory, including the American
Mathematics Society, usually make the claim that there is no perfect voting
system. The equivalent statement in thermodynamics is that there is no
perpetual motion machine.
As you point out, that does not preclude voting methods of different
efficiency, the equivalent of heat engines of differing efficiency. The
engines depend on efficient transfer of surplus heat, to work requirements,
to keep the engine going. Similarly, transfers of vote surpluses, to
elective quotas, keep the count procedure going. Heat forms a random
distribution of motion. And votes typically form a random distribution of
choice (subject to left or right skews).
Binomial STV would perhaps be rather more efficient than traditional STV,
because it rationally conserves exclusion information. In rough analogy, a
binomial STV “heat engine” is better “insulated,” to conserve heat.
Thermodynamics is not just a dynamic of heat but also its insulation, in a
closed system. Likewise, an election method is not just an active election,
but also a closed system of exclusion.
Regards,
Richard Lung.
Election-Methods mailing list - see https://electorama.com/em for list
info
On 05.06.2022 19:16, Carl Schroedl wrote:
As a software guy, the connection I make is to something I have wondered
for a while -- whether it is useful to study social choice functions as
lossy compression algorithms. I haven't thought it through, but it could
be interesting to see if the rate-distortion branch of information
theory would apply.
If you're trying to design a method that has the best possible VSE for a
ranked voting method, then it may be possible to use ideas from vector
quantization. In a spatial model, the voters rank the candidates
according to proximity, and then the method finds the winner that's
closest to the voters using only this information. So it's trying to
find a vector (n-dimensional point) that's closest, in an Euclidean
sense, to the distribution of the voters... though unlike ordinary VQ,
it doesn't know the actual distances, only their ranking.
Similarly, I'd say quota-based proportional representation is like
clustering. Monroe's method is the most obvious clustering-like PR
method: you assign each candidate a voter, so that each candidate has
the same number of voters, and so that the total voter-candidate
distance is minimized. (One possible objection to Monroe is that it
doesn't care about what the voter thinks about the composition of the
rest of the assembly, just his preferred candidate.)
That's for honest voters, though. With strategic voting, the
"compression method" (clustering method) becomes partly adversarial:
keep the outcome from degrading too much if some fraction of the votes
is arbitrarily altered.
-km
You mention sincere vs insincere voting ... which leads to game theory. In
game theory most optimal strategies are mixed ... stochastic combinations
of pure (deterministic) strategies.
So it is a matter of luck if it turns out that a deterministic strategy is
optimal.
We think of Approval as a deterministic method, but that's only because we
have externalized optimal strategy considerations to the (cagey) voters and
their (mostly gut level) probability estimates.
Back to multi-winner methods. A rule of thumb for a minimum number of seats
for good proportional representation is the reciprocal of S=Sum (p_i)^2,
where p_i is the probability that candidate i would get elected by random
favorite ballot.
In general, Sum p_i*r_i is a weighted arithmetic mean of the r values,
where the p values are the normalized weights.
So the given sum S is a kind of mean value of the p values. If there were n
of them, and they were all equal, the mean would be 1/n, so that the rule
of thumb would yield 1/(1/n), that is n, which makes perfect sense.
The same would work for any other kind of weighted mean. For example the
weighted geometric mean:
G=Prod(r_i^p_i) is a weighted geometric mean of the r values where the p
values are the (,normalized) weights.
If the r vector is a copy of the p vector we get
G=Prod(p_i^p_i)
If we take the log of the reciprocal of G, we get ...
Log(1/G)=-log(Prod(p_i)^p_i), which expands to -Sum(p_i*log p_i), which we
recognize as the Shannon Information/ entropy formula.
A local global max of this entropy occurs when the distribution is uniform,
that is when p_i=1/n.
So it turns out that the rule of thumb formula n=1/S is related to the
Shannon information/Entropy of the favorite candidate lottery
distribution. In fact, the log of the rule of thumb value is a good
approximation to the Shannon information ... that is
log(1/S) ~ log(1/G), in general, and the approximation straightens out to
equality if all the p values are equal, or if all are zero except one.
So we begin to see connections between statistical mechanics and the
various distributions that are so ubiquitous in voting methods. These
distributions include mixed strategy distributions, distributions of voters
and candidates in various issue spaces, etc.
Let's keep our eyes open for more connections!
-Forest
El dom., 5 de jun. de 2022 10:44 a. m., Kristofer Munsterhjelm <
km_elmet@t-online.de> escribió:
On 05.06.2022 19:16, Carl Schroedl wrote:
As a software guy, the connection I make is to something I have wondered
for a while -- whether it is useful to study social choice functions as
lossy compression algorithms. I haven't thought it through, but it could
be interesting to see if the rate-distortion branch of information
theory would apply.
If you're trying to design a method that has the best possible VSE for a
ranked voting method, then it may be possible to use ideas from vector
quantization. In a spatial model, the voters rank the candidates
according to proximity, and then the method finds the winner that's
closest to the voters using only this information. So it's trying to
find a vector (n-dimensional point) that's closest, in an Euclidean
sense, to the distribution of the voters... though unlike ordinary VQ,
it doesn't know the actual distances, only their ranking.
Similarly, I'd say quota-based proportional representation is like
clustering. Monroe's method is the most obvious clustering-like PR
method: you assign each candidate a voter, so that each candidate has
the same number of voters, and so that the total voter-candidate
distance is minimized. (One possible objection to Monroe is that it
doesn't care about what the voter thinks about the composition of the
rest of the assembly, just his preferred candidate.)
That's for honest voters, though. With strategic voting, the
"compression method" (clustering method) becomes partly adversarial:
keep the outcome from degrading too much if some fraction of the votes
is arbitrarily altered.
Election-Methods mailing list - see https://electorama.com/em for list
info
El dom., 5 de jun. de 2022 3:06 p. m., Forest Simmons <
forest.simmons21@gmail.com> escribió:
You mention sincere vs insincere voting ... which leads to game theory. In
game theory most optimal strategies are mixed ... stochastic combinations
of pure (deterministic) strategies.
So it is a matter of luck if it turns out that a deterministic strategy is
optimal.
We think of Approval as a deterministic method, but that's only because we
have externalized optimal strategy considerations to the (cagey) voters and
their (mostly gut level) probability estimates.
Back to multi-winner methods. A rule of thumb for a minimum number of
seats for good proportional representation is the reciprocal of S=Sum
(p_i)^2, where p_i is the probability that candidate i would get elected by
random favorite ballot.
In general, Sum p_i*r_i is a weighted arithmetic mean of the r values,
where the p values are the normalized weights.
So the given sum S is a kind of mean value of the p values. If there were
n of them, and they were all equal, the mean would be 1/n, so that the rule
of thumb would yield 1/(1/n), that is n, which makes perfect sense.
The same would work for any other kind of weighted mean. For example the
weighted geometric mean:
G=Prod(r_i^p_i) is a weighted geometric mean of the r values where the p
values are the (,normalized) weights.
If the r vector is a copy of the p vector we get
G=Prod(p_i^p_i)
If we take the log of the reciprocal of G, we get ...
Log(1/G)=-log(Prod(p_i)^p_i), which expands to -Sum(p_i*log p_i), which
we recognize as the Shannon Information/ entropy formula.
A local global max of this entropy occurs when the distribution is
uniform, that is when p_i=1/n.
To round out this part of the discussion I should have pointed out that the
global min of entropy is zero, which occurs onlywhen all values (except
one) of p are zero, corresponding to a single winner (n=1) election in this
context.
In the strategy context it would signify a pure/deterministic (as opposed
to mixed) optimal strategy.
So it turns out that the rule of thumb formula n=1/S is related to the
Shannon information/Entropy of the favorite candidate lottery
distribution. In fact, the log of the rule of thumb value is a good
approximation to the Shannon information ... that is
log(1/S) ~ log(1/G), in general, and the approximation straightens out to
equality if all the p values are equal, or if all are zero except one.
So we begin to see connections between statistical mechanics and the
various distributions that are so ubiquitous in voting methods. These
distributions include mixed strategy distributions, distributions of voters
and candidates in various issue spaces, etc.
Let's keep our eyes open for more connections!
-Forest
El dom., 5 de jun. de 2022 10:44 a. m., Kristofer Munsterhjelm <
km_elmet@t-online.de> escribió:
On 05.06.2022 19:16, Carl Schroedl wrote:
As a software guy, the connection I make is to something I have wondered
for a while -- whether it is useful to study social choice functions as
lossy compression algorithms. I haven't thought it through, but it could
be interesting to see if the rate-distortion branch of information
theory would apply.
If you're trying to design a method that has the best possible VSE for a
ranked voting method, then it may be possible to use ideas from vector
quantization. In a spatial model, the voters rank the candidates
according to proximity, and then the method finds the winner that's
closest to the voters using only this information. So it's trying to
find a vector (n-dimensional point) that's closest, in an Euclidean
sense, to the distribution of the voters... though unlike ordinary VQ,
it doesn't know the actual distances, only their ranking.
Similarly, I'd say quota-based proportional representation is like
clustering. Monroe's method is the most obvious clustering-like PR
method: you assign each candidate a voter, so that each candidate has
the same number of voters, and so that the total voter-candidate
distance is minimized. (One possible objection to Monroe is that it
doesn't care about what the voter thinks about the composition of the
rest of the assembly, just his preferred candidate.)
That's for honest voters, though. With strategic voting, the
"compression method" (clustering method) becomes partly adversarial:
keep the outcome from degrading too much if some fraction of the votes
is arbitrarily altered.
Election-Methods mailing list - see https://electorama.com/em for list
info
On 06.06.2022 00:06, Forest Simmons wrote:
You mention sincere vs insincere voting ... which leads to game theory.
In game theory most optimal strategies are mixed ... stochastic
combinations of pure (deterministic) strategies.
So it is a matter of luck if it turns out that a deterministic strategy
is optimal.
Yes. I'm just referring to that physics usually doesn't "fight back" the
way strategic voters do :-)
We think of Approval as a deterministic method, but that's only because
we have externalized optimal strategy considerations to the (cagey)
voters and their (mostly gut level) probability estimates.
Back to multi-winner methods. A rule of thumb for a minimum number of
seats for good proportional representation is the reciprocal of S=Sum
(p_i)^2, where p_i is the probability that candidate i would get elected
by random favorite ballot.
A related connection is to Laakso and Taagepera's effective number of
parties: https://electowiki.org/wiki/Effective_number_of_parties which
describes a party distribution as equivalent to a certain number of
"equally sized" parties. E.g. a dominant-party system may have an ENP
measure of 1.3, signifying that the second party has significantly less
support (or representation) than the first.
This measure is also 1/p_i^2 and has been criticized for being not
uniform enough (i.e. not providing the same results for lots of small
parties as a few large ones). Greene and Bevan suggested an information
entropy measure instead, as I referenced in the Electowiki article.
In addition, for party list PR, Webster's method minimizes the
Sainte-Laguë index of disproportionality, sum over parties i: (seats
given to party i - votes obtained by i)^2/(votes obtained by i). This
is, if I recall correctly, related to the chi-squared test statistic:
x^2 = sum over categories i: (observed counts for i - expected counts
for i)^2/(expected counts for i)
The chi-squared statistic in turn is an approximate G-test. The G-test's
form is:
G = 2 sum over categories i: (observed counts for i) * ln ( observed_i /
expected_i),
which looks vaguely like an entropy term. Wikipedia says it's related to
mutual information, but I know too little about this to comment.
https://en.wikipedia.org/wiki/G-test#Relation_to_mutual_information
(Finally, speaking of random favorite, I wrote a post about a
semiproportional determinization of it for multiwinner, here:
http://lists.electorama.com/pipermail/election-methods-electorama.com/2019-October/002323.html
Its proportionality in the limit, yet possible bad results with only a
few seats shows the limits of PR as lotteries, I think. At least for the
random favorite lottery. But perhaps the idea of eliminating the winner
can be used to extend Plurality-based PR to ranked method PR... or be
used as a component of a Condorcetian multiwinner method.)
-km
El dom., 5 de jun. de 2022 4:05 p. m., Kristofer Munsterhjelm <
km_elmet@t-online.de> escribió:
On 06.06.2022 00:06, Forest Simmons wrote:
You mention sincere vs insincere voting ... which leads to game theory.
In game theory most optimal strategies are mixed ... stochastic
combinations of pure (deterministic) strategies.
So it is a matter of luck if it turns out that a deterministic strategy
is optimal.
Yes. I'm just referring to that physics usually doesn't "fight back" the
way strategic voters do :-)
That's what I thought; I was just using it as a segue into another use of
probability in voting theory in the general theme of Robert
Bristow-Johnson's question.
You are way ahead of me!
We think of Approval as a deterministic method, but that's only because
we have externalized optimal strategy considerations to the (cagey)
voters and their (mostly gut level) probability estimates.
Back to multi-winner methods. A rule of thumb for a minimum number of
seats for good proportional representation is the reciprocal of S=Sum
(p_i)^2, where p_i is the probability that candidate i would get elected
by random favorite ballot.
A related connection is to Laakso and Taagepera's effective number of
parties: https://electowiki.org/wiki/Effective_number_of_parties which
describes a party distribution as equivalent to a certain number of
"equally sized" parties. E.g. a dominant-party system may have an ENP
measure of 1.3, signifying that the second party has significantly less
support (or representation) than the first.
This measure is also 1/p_i^2 and has been criticized for being not
uniform enough (i.e. not providing the same results for lots of small
parties as a few large ones). Greene and Bevan suggested an information
entropy measure instead, as I referenced in the Electowiki article.
In addition, for party list PR, Webster's method minimizes the
Sainte-Laguë index of disproportionality, sum over parties i: (seats
given to party i - votes obtained by i)^2/(votes obtained by i). This
is, if I recall correctly, related to the chi-squared test statistic:
x^2 = sum over categories i: (observed counts for i - expected counts
for i)^2/(expected counts for i)
The chi-squared statistic in turn is an approximate G-test. The G-test's
form is:
G = 2 sum over categories i: (observed counts for i) * ln ( observed_i /
expected_i),
which looks vaguely like an entropy term. Wikipedia says it's related to
mutual information, but I know too little about this to comment.
https://en.wikipedia.org/wiki/G-test#Relation_to_mutual_information
(Finally, speaking of random favorite, I wrote a post about a
semiproportional determinization of it for multiwinner, here:
http://lists.electorama.com/pipermail/election-methods-electorama.com/2019-October/002323.html
Its proportionality in the limit, yet possible bad results with only a
few seats shows the limits of PR as lotteries, I think. At least for the
random favorite lottery. But perhaps the idea of eliminating the winner
can be used to extend Plurality-based PR to ranked method PR... or be
used as a component of a Condorcetian multiwinner method.)
-km