Expanding Elo to Evaluate Individual Player Performance
An approach to objectively rate football players using an Elo-based algorithm adapted for team sports.
The following summary critically reviews the research paper titled “A football player rating system” by Stephan Wolf, Maximilian Schmitt and Björn Schuller. All data, figures, and analysis presented here are drawn from their original work; I do not claim any authorship or ownership of the content. This summary has been written to provide a concise and technically informed synthesis of the paper’s findings, methodologies, and implications, while maintaining fidelity to the authors’ intellectual contributions.
1. Introduction
Traditional player rating systems in football generally fall into two categories: subjective evaluations and objective statistics-based approaches. In subjective systems, experts or journalists assign ratings based on their perception of performance. In objective systems, ratings are derived from recorded match data such as passes, duels, or other events, often using statistical or machine learning techniques. Both approaches have limitations. Subjective ratings can be influenced by bias, while event-based models often require extensive tracking or event data that are not universally available.
The paper “A football player rating system” introduces a new approach: an adaptation of the Elo algorithm, originally developed for individual sports such as chess, to evaluate individual players in football. The key idea is to construct an objective and adaptive rating system that relies only on official match reports. These reports include the final score, lineups, substitutions, and minutes played.
The authors argue that football is a team sport in which individual performance cannot be fully separated from team performance. Therefore, instead of relying on detailed event data, the proposed system evaluates players based on match results while accounting for the strength of teammates and opponents. If a player achieves better results than expected, given the strength of both teams, their rating increases. If the result falls short of expectations, their rating decreases.
Figure 1 in the paper illustrates the core idea: a player’s previous rating is compared to an expectation derived from their rating and the opposing team’s rating. The difference between expectation and actual result determines the rating update. The more unexpected the result, the larger the rating change.

2. The Elo Algorithm
In various sports, Elo rating systems, which adapt the rating of players based on their results and the relative skill level of their opponents, have been well established. The American statistician Arpad Emrick Elo developed the Elo algorithm in 1960, which has since been used by the World Chess Federation (FIDE) for calculating FIDE ratings; objective scores indicating the playing strength of chess players. This system has stood the test of time as a highly effective method for evaluating performance, evolving beyond chess and finding applications in sports like tennis, esports, and even team-based games.
The original Elo algorithm calculates player ratings based on their performance relative to their opponents. It is particularly effective in sports like chess and tennis, where matches are between two players. In these head-to-head matchups, Elo ratings are adjusted depending on the outcome of the match, with winners gaining points and losers losing points. The change in rating is determined by the expected outcome—underdogs gain more points if they win, while favorites gain fewer. This study adapts the Elo algorithm for football, accounting for the complexity of a team-based sport where multiple players influence the outcome.
Applying the Elo algorithm to a team sport like football requires accounting for additional variables such as team interactions, substitutions, and the influence of match locations. Unlike individual sports, where each player’s influence is more direct, football involves shared responsibilities, making it imperative to devise a system that attributes performance accurately across players who participate in the match.
3. Football Player Rating System
The Elo-based football player rating system adapts the traditional Elo model to assess individual players in a team environment. Key features and modifications include:
3.1 Algorithm
The Elo-based system for football was designed to evaluate individual players in a team context. The traditional Elo algorithm was modified to consider match outcomes, individual contributions, and team-level factors. This adapted model uses match data to adjust player ratings after each game, taking into account factors such as goal differential, minutes played, and home advantage. The model also distinguishes between starting players and substitutes, recognizing that players who start a match and those who come on later may have different impacts depending on match conditions and fatigue levels.
The home advantage is explicitly factored into the ratings, as playing at home has been statistically shown to benefit teams across leagues.
After the match, each player’s result is determined based on the goal difference while they were on the pitch. If their team scored more goals than it conceded during their minutes played, this is treated as a win for that player. If goal difference was zero, it is treated as a draw. If negative, it is a loss.
The magnitude of the rating change depends on:
whether expectations were exceeded or not
the clarity of the result (goal difference)
the number of minutes played
Clearer results (for example, larger goal differences) lead to larger rating adjustments. For draws, the number of minutes played affects the magnitude of the change.
Importantly, the update combines two components:
the individual player’s result
the overall team result.
This reflects the idea that players influence team outcomes, but team outcomes also influence individual evaluations.
Moreover, the algorithm takes into account contextual factors like whether a team is playing against a significantly stronger or weaker opponent. Players who perform well against stronger opponents receive a greater boost to their rating, reflecting the difficulty and importance of their contributions under challenging conditions. This adjustment ensures that players who excel in tough matches are appropriately rewarded, giving a more nuanced view of their abilities compared to players who only perform well against weaker teams.
These values evolve over time. If a player participates regularly, their rating becomes more stable. If a player does not play for a while, their rating becomes more sensitive to new results. When a player changes clubs, both parameters are reset to allow faster adaptation to the new team context.
The team itself also receives an updated rating after each match using a similar mechanism.
3.2 Data
The system was applied to 17,086 matches from 18 European competitions between the 2014/15 and 2017/18 seasons. These include:
16 domestic leagues (such as the Premier League, La Liga, Serie A, Bundesliga, etc.)
the UEFA Champions League
the UEFA Europa League
All match data was collected from the website kicker.de. Only information contained in official match reports was used.
In total, 11,139 different players and 438 clubs were registered in the system.
Because the model requires only lineups, substitutions, goals, and minutes played, it can be applied broadly without relying on detailed event or tracking data.
3.3 Player Initialization
To effectively rate players, an initial rating must be assigned. Players who lacked prior professional experience were initialized based on their team’s average rating or their teammates’ ratings. This approach ensures that new players start with a realistic rating that reflects their team environment. Young players, especially those making their debut in strong teams, may initially receive ratings that do not fully capture their potential. Future adjustments to add a youth-specific bonus could improve the accuracy of these initial ratings.
In the current system, new players joining from lower leagues or academies are given provisional ratings, which can change significantly over their first few matches. Players who perform well in their debut matches are rapidly adjusted upwards, allowing the system to quickly reflect their actual skill level. This dynamic initialization process helps in providing a fair starting point while allowing for quick corrections based on on-field performance. This approach is particularly useful for capturing the trajectory of emerging talents, whose performance can improve rapidly as they gain experience and confidence.
Another aspect of player initialization is accounting for players returning from injuries. When a player returns after an extended absence, their rating may be adjusted to reflect the uncertainty around their current form. This prevents players from being unfairly penalized or overly rewarded based on outdated performances. Instead, their rating is allowed to stabilize over a series of matches, giving a more accurate reflection of their current abilities.
The system also tracks players who transfer between leagues with different competitive levels. Players who move from weaker leagues to stronger ones may start with a rating adjustment to reflect the higher level of competition, while those moving in the opposite direction might see their initial ratings decrease. This adjustment ensures a smoother transition and a more realistic comparison between players from different competitive environments.
3.4 Player Impact
Because player ratings are strongly linked to team performance, the authors introduce an additional metric called player impact.
Player impact measures how a team’s rating changes when the player is on the pitch, compared to when the player is not on the pitch.
This metric is calculated using team rating changes per minute, with results weighted more heavily for recent matches (using a half-life of one year).
Player impact helps identify players whose presence is associated with stronger team performance, even if their overall rating is influenced by team quality.
4. Results
4.1 Presentation of Results
The system successfully generated ratings for individual players and teams, revealing notable insights:
Top Players: Gerard Piqué, Lionel Messi, and David Alaba emerged as the highest-rated players based on performance metrics, with ratings exceeding 4900 points.

Players like Mohamed Salah and Kevin De Bruyne ranked highest in terms of player impact, which measures the influence a player has on their team’s performance. Lionel Messi’s contributions were not only reflected in goal-scoring but also in his role in creating opportunities for teammates. The ratings showcased the value of players who contribute significantly to both offense and defense, providing a balanced view of a player’s overall importance.

4.2 Evaluation of the Results
The Elo-based player rating system demonstrated slightly better predictive accuracy for match outcomes compared to a team-only rating approach. To evaluate this, the authors compared two methods:
A team rating method, which evaluates the strength of a team regardless of the specific lineup.
A player rating method, which calculates team strength based on the average rating of the players in the lineup for that specific match.
In both approaches, the prediction rule was simple: the team with the higher rating was predicted to win.
For the full dataset across four seasons, the player-based method correctly predicted 52.55% of matches, compared to 52.29% for the team-based method. Although the difference is modest, it consistently favors the player-level approach.

The improvement becomes clearer over time. Because players are initially rated based on their team, the two methods produce identical predictions at the beginning of the evaluation period. As more matches are processed and player ratings diverge from team averages, the player-based approach begins to outperform the team-based one.

The authors also evaluate predictive performance using the Brier score, which measures the average squared difference between the predicted expectation and the actual match result. A lower Brier score indicates better predictive accuracy.

For the 2017/18 season across major leagues and the Champions League, the player-based method consistently achieved lower Brier scores than the team-only approach and a simple baseline model that assigns equal strength to both teams.

Overall, the evaluation shows that incorporating individual player ratings into match strength calculations leads to slightly more accurate match predictions. The improvement is not dramatic, but it is systematic across competitions and seasons. This supports the authors’ central claim that knowing the ratings of the players in the lineup provides additional predictive value beyond using team ratings alone.
5. Potentials and Limitations of the Rating System
While the system shows potential for objectively comparing players, several limitations and areas for future development exist.
5.1 Youth Player Initialization
One limitation concerns the initialization of young players. Since new players are typically initialized at the level of their teammates, this may not fully reflect the developmental trajectory of young talents. The authors suggest that introducing a “youth bonus” mechanism could improve this aspect. For example, younger players could initially receive a reduced rating that adjusts upward as they accumulate match experience.
Such an adjustment would allow the model to better account for development over time.
5.2 Amateur Level Application
A major strength of the system is that it requires only information contained in match reports. Because official match reports exist not only in professional football but also in many amateur leagues, the system could potentially be extended to evaluate amateur players.
The authors propose heuristic formulas for initializing amateur teams based on league level. This would allow the rating system to scale beyond professional football and compare players across different divisions and regions.
This scalability is one of the system’s key advantages compared to approaches that require detailed event or tracking data.
5.3 Player Transfers
When a player changes clubs, the system resets the adaptive parameters (k- and q-values) to allow the rating to adjust more quickly to the new team environment. This ensures that a player’s rating can respond dynamically to changes in competitive context.
The authors note that while the transfer mechanism allows adaptation, football remains a sport influenced by randomness and contextual effects. Therefore, no rating system can perfectly predict future performance.
6. Conclusion
The Elo-based football player rating system provides a novel, objective, and adaptive approach to evaluating football players. Unlike purely subjective or statistics-based methods, this model incorporates team-level success, individual contributions, and adjustments for factors such as home advantage and player transfers.
Applied to more than 17,000 matches and over 11,000 players across European competitions, the system demonstrates scalability and slightly improved predictive accuracy compared to team-only rating approaches.
The predictive capabilities of the Elo-based system also suggest applications beyond team evaluation, such as betting markets or fantasy football. These markets could benefit from the system’s ability to objectively assess individual player performance, offering users a more data-driven way to make informed decisions. Moreover, coaches and analysts could use the model to evaluate the impact of tactical changes and player rotations, ultimately improving team strategies.
The authors conclude that the method represents an innovation compared to many existing models, particularly because it requires only widely available match reports. They also highlight future research directions, including possible extensions of Elo models that incorporate additional variables such as time or margin of victory.
Overall, the system offers a practical framework for ranking players, comparing teams, and improving match prediction while maintaining simplicity and broad applicability.
Learn More
Here you have a direct link to get the free xG Mini-Guide, suited for those who want to fully understand the model in 5 minutes.
References
Wolf, S., Schmitt, M., & Schuller, B. (2021). A football player rating system. Journal of Sports Analytics, 6(4), 243-257. https://content.iospress.com/articles/journal-of-sports-analytics/jsa200411









Fascinating article! The part that got me the most excited is the amateur player evaluation... feels like teams are losing a goldmine because of the data blackout that occurs at the deepest levels