Pages

Showing posts with label neural. Show all posts
Showing posts with label neural. Show all posts

Friday, April 22, 2016

Part 6: Review of Game 1: Lee Sedol underestimates AlphaGo's incredible fighting power (The historic match of deep learning AlphaGo vs. Lee Sedol)


[NL Versie]

Review of Game 1: Lee Sedol underestimates AlphaGo's incredible fighting power


The first game of the Google DeepMind challenging match  between deep learning AlphaGo and top Go professional Lee Sedol (9p) is the most important game of this match as this is the first time an AI program plays without handicap against one of the strongest human Go-players of the world.


AlphaGo has not learned from professional games and therefore, as repeatedly stated and emphasized by Demis Hassabis, has not a single game of Lee Sedol in it's database. Lee Sedol doesn't know much more about the program than what he has seen from AlphaGo's match against Fan Hui (2p) last October: a strong but now and then malfunctioning program, sometimes making obvious mistakes in complex situations, without too much whole-board understanding. And that against a substantially weaker Go-player who made many overplays Lee Sedol probably never would make. 

Both sides will thus be scanning each other for the very first time, finding out the strategies, way of thinking, patterns, playing strength of the opponent and exploit any possible weaknesses. While AlphaGo originally trained on 130,000 amateur games (up to 8-9 dan = ~1-2p) and further improved from millions of self-play games, Lee Sedol played several thousands of top-tournament games in the almost 25 years of his professional Go-life. 





The opening moves of the first game in this historic match between mankind and AI are played while it is early in the morning 5 AM in the Netherlands. Here the match is watched under extremely high-voltage excitement in the European Go Cultural Center (EGCC) by more than 60 people that have come from all corners of The Netherlands with or without sleeping bag, to experience this once-in-a-lifetime event together real-time and to live through the game as conscious as possible. 

Two big screens are placed on the wall from which the online game commentary and playing room of the match are displayed. Huge posters, a demonstration board and many laptops with all-sides discussions of the match worldwide. Meanwhile the game is discussed by the strongest European players we have: Merlijn Kuin (6d), Peter Brouwer (6d), and Guo Juan (5p, inactive as pro player).

It is no secret that Lee Sedol consulted several Go-players and programmers in preparation to this match. This may be also one of the reasons why Lee Sedol will test in specific manners how AlphaGo will treat and react on moves that the program never has seen (or learned) before. Besides, Lee Sedol is assisted by a team of psychologists and coaches to deal with the enormous mental pressure and media attention worldwide. And to let himself not being brought out-of-balance by the non-ability to assess the emotional state and mental stability of his opponent (in this case, the biggest challenge is actually to win from himself in particular). 


After Lee Sedol's move with black 7 (see Dia. 1) Merlijn Kuin (6d) kicks off immediately: “this is a rather unique move and very aggressive of black. According to me, this is a position which appeared never before on the board. I am convinced that Lee Sedol uses a  preconceived strategy that fits his style". Lee Sedol makes a bad decision by choosing an unusual opening of which he is sure it is absent in AlphaGo's database. However, Lee Sedol himself is also unfamiliar with such a rare and less optimal opening pattern. 


Dia. 1:  Game 1, after white 14 (lowest triangle, Lee Sedol is black).
Black's move 7 is marked with a green dot. 
When seeing the first moves played by AlphaGo, Guo Juan (5p) shakes her head in a clearly disapproving manner (Dia. 1): "this program really needs a good teacher who will knock and bestow upon the head when playing such a bad move, every just knows that this move isn't good at all ... these first white moves are obviously not optimal and look like beginner's mistakes". The order in which the white moves are played (marked with a triangle in Dia. 1), provide black with the opportunity to answer as in the game. Therefore, the result is somewhat less for AlphaGo (see Dia. 1).


Dia. 2: 
 Game 1, after black 23 (circle, Lee Sedol is black) 
With Lee Sedol's move 23 (circle in Dia. 2) the game is going up in flames immediately. Guo Juan (5p) comments almost instantly: "this move by Lee Sedol is much too compulsory and forcing" and she indicates that black's tsuke (touching move) is a huge overplay. The tone of the game has been set as the rest of the game flow will be determined by Lee Sedol's huge underestimation of the program. AlphaGo cuts at once and takes the initiative by putting Lee Sedol under high pressure.

After barely 25 moves (Dia. 3), there are seven groups fighting each other. Peter Brouwer (6d) remarks that the shapes and patterns in this game are not particularly beautiful. And Michael Redmond (9p) comments: "in reaction to a slightly less optimal move by AlphaGo, it appears that Lee Sedol comes at once with an overplay. This human reaction is well understandable and makes AlphaGo at least one stone stronger". 


Dia. 3:  Game 1, after white 50 (circle, Lee Sedol is black) 
With move 48 (triangle in Dia. 3), white attacks the black group in the upper right-hand side. Black answers (square in Dia. 3) and adds some strength to his group by creating additional eye potential. In case of emergency, black can connect underneath with the three black stones at the top right side (e.g. by black R18). However, since it is far too early for black to play himself on 48 (triangle in Dia. 3), white's move 48 is seen by top profs as a mistake by AlphaGo (even though it is sente).


Dia. 4:  Game 1, after white 80 (white stone with square)  
After AlphaGo pushes from behind on the seventh line in Dia. 4, it becomes clear that the program does not respect widely accepted Go-proverbs ('in the far east a child gets a smacking or a hit with a japanese fan on his hands"). 

With AlphaGo's strategy Lee Sedol gets huge influence in the center. But at the expense of a much weaker corner at the bottom right. After black neutralizes most of the aji of the two white cutting stones in the center (triangle in Dia. 4), white strikes and invades  (circle in Dia. 4). Evidently, AlphaGo wants to use the force of the white wall and to transform that into fighting power.

Lee Sedol plays tenuki and attacks the white stone in the bottom left corner while expanding his center moyo (blakc stone with square in Dia. 4). In turn, AlphaGo plays tenuki as well and captures two black stones, thereby removing most dangers (aji) around the center position (white stone with square in Dia. 4).

However, since removing black's aji is not urgent right now and with this move AlphaGo makes an important error in positional judgement. After the program's blunder, top prof Gu Li (9p) stated: "Lee has now a chance of 90% to win this game". And Michael Redmond comments: "AlphaGo is very accurate and profound in computing moves and very balanced in influence and territory. After an earlier better position for AlphaGo, Lee Sedol has brought the game in balance again".


Dia. 5: 
 Game 1, after white 102 (circle, Lee Sedol is black)
And then suddenly in a calmly developing game, a dramatic upheaval happens: "AlphaGo's move 102 (Dia. 5) really is a superhuman move" according to Michael Redmond. It looks like white has carefully prepared this invasion with moves in the upper right and with the corner approach --25 moves earlier-- at the bottom right (white 78, circle in Dia. 4). 

A top Go-prof who thoughtfully suggested this move was laughed at in his face: initially many Go-profs are rather negative and demand proof for this move. AlphaGo just plays it as nothing special. In second thought, however, profs worldwide admire this wonderful effective invasion. Others are heavily  disconcerted by this remarkable move of the AI program as it shows it's incredible strength and profound 'understanding' of the concepts the game (including aji). This clearly shows how different this version of AlphaGo is compared to that during the Fan Hui match. 

With this invasion AlphaGo attempts to blow up entirely Lee Sedol's position at the right side of the board in a terrifying fight (circle Dia. 5) while using the great influence it has built earlier. At the same time, this invasion gives opportunities to make a base for white's group by exploiting the weaknesses in black's position on the right-hand side (note how the white group extends over the full length of the board, see Dia. 6).


Dia. 6:  Game 1, after white 116 (circle, Lee Sedol is black)
With AlphaGo's impressively strong and fabulous move, Lee Sedol has no choice other than to give up is three stones in the upper right (see Dia. 6). In exchange, he gets three white stones but black's potential at the right side has disappeared like frost under the morning sun (in gote). Then, AlphaGo plays a rather slow (and a bit cowardly as it seems unnecessary right now and is not optimal according to the commentators) but surely effective move which secures the white corner in the upper left (circle in Dia. 6).

To play this move right now in the game (while there are several other, larger moves) indicates that the program believes it is ahead and has taken a sufficient lead to win the game. AlphaGo is absolutely not interested in playing the most optimal, largest or most efficient moves: the program only plays moves that provide the highest probability of winning the game.


Dia. 7:  Game 1, after black 127 (circle, Lee Sedol is black)
Anyhow, Lee Sedol starts a fight in the bottom right corner in an attempt to put AlphaGo under pressure (Dia. 7). Due to some mistakes by black,  white succeeds in making a living group that is much larger than what black should have allowed. Lee Sedol seems already too much behind to catch up in the rest of the game.


Dia. 8:  Game 1, after white 136 (circle, Lee Sedol is black) 
Nonetheless, Lee Sedol checks AlphaGo's whole-board awareness (triangle in Dia. 8): does white notice that his long, elongated group is not yet alive? With the effective defense by AlphaGo it becomes clear that the program has excellent whole-board understanding (circle in Dia. 8).


Dia. 9: Game 1, final position after white 186 (circle, Lee Sedol is black)
After about 50 endgame moves, Lee Sedol resigns. With AlphaGo's move 186 (circle in Dia. 9), Lee Sedol is more than 5 points behind. No doubt, Lee Sedol is greatly saddened by his severe underestimation of the mind-blowing strength and amazing fighting of this version of AlphaGo (as opposed to that during the Fan Hui match about half year ago).





This is the first time in history that an AI program defeats a top Go-prof in a formal game without handicap. Go-profs worldwide are speechless and all agree that this has been a fabulous and historical game. And of course, there is a lot more to say in detail about the firm attacks, remarkable, aggressive and deliberate moves, tactics and strategies of both sides in this game. 

The era in which a human could wipe an AI program (without handicap) from the board is past and once and for all closed. This magnificent and superbly profound way of playing Go is never shown before on our planet. And will leave behind incredibly deep imprints and heavy tracks on the hundreds of millions of people worldwide who were following this game online.


Lee Sedol loses this first game against AlphaGo and reacts deeply saddened, disappointed and very touched: "I was very surprised because I did not think that I would lose the game. I am shocked by how well AlphaGo played, I can admit that. I didn't think that AlphaGo would play the game in such a perfect manner. A notable mistake from my side at the beginning of the game kept continuing during the game and lasted until the very last. AlphaGo's early strategy was excellent and I was stunned by one unconventional move Alphago made (see Dia. 5) that a human never would have played. I am in shock, but what's done is done".

Despite his loss in this opening game, Lee Sedol stated he did not regret accepting the challenge: "I had a lot of fun playing Go this game and I'm looking forward to the future games". 


Michael Redmond (9p): "In this first game of the series, AlphaGo triumphed by a very narrow margin. While Lee Sedol had led for most of the match, AlphaGo managed to build up a strong lead in the closing stages of the game. AlphaGo played profound and solid moves with which it secured victory and ultimately has punished Lee Sedol severely and consequent for his lesser opening move (move 7) and aggressive overplay (move 23)". 


In this first game, AlphaGo has shown evidently  that it has a taste for aggressive and offensive play, something the program not necessarily demonstrated during the Fan Hui match. In this game AlphaGo has exhibited surprisingly accurate, profound, and versatile judging of the position.

And AlphaGo has played at least one incredibly wonderful move that achieved so many goals at the same time that it is hard to apprehend (Dia. 5). During the fight that followed, the program succeeded in getting enough advantage so that it was confident to secure the upper left corner with an all-telling move and claiming the victory of the game (Dia. 6).



Although Lee Sedol controlled the majority of the game, AlphaGo managed to rebuild itself in a phenomenal way and secured a clear advantage during the last 20 minutes of the game after which Lee Sedol resigned.

With all his hubris and self-assurance prior to this match, his underestimation of the startling, mind-boggling, fabulous, and wonderful play and strength of AlphaGo, his expectations on the basis of the games during the Fan Hui match, his immense concentration and involvement with this first game of the match, and the enormous media attention and psychological pressure by all eyes of the world, this extremely surprising and unexpected loss must have felt particular anguish for Lee Sedol.

Friday, April 8, 2016

Part 4: Artificial Intelligence versus Lee Sedol: Expectations and Predictions (The historic match of deep learning AlphaGo vs. Lee Sedol)

AI vs. Human: Expectations and Predictions of the Match

Among the strongest Go players of the world there are almost none who really believe that deep learning AlphaGo will be able to defeat Lee Sedol in the Google DeepMind challenging match. However, many of them are eager to play against AlphaGo themselves. They expect this based on the games of AlphaGo against Fan Hui, the apparent mistakes AlphaGo made from time to time, and the fact that Fan Hui made a number of mistakes that Lee Sedol is very unlikely to make. A few examples of what top 9p players have been saying in advance about the outcome of the match: 

Changho Lee (9p): “I heard about the match between Lee Sedol and AlphaGo. I am surprised that an AI program could challenge a human pro on even game. I believe DeepMind challenged Lee Sedol because they thought AlphaGo has chance of winning. It will be an interesting match, but I think Lee Sedol will win this time”. 

Ke Jie( 9p): “I used to think AI can never beat human, at least it won't happen within 10 years. But this unbelievable .. I think Lee Sedol will win the match in March”.   

Dongyoon Kang (9p): “I went through the game record of the match between AlphaGo and Fan Hui, and AlphaGo plays really well. It makes huge mistakes where standard procedures are necessary. I do not understand how computers can make such mistakes.  I think AlphaGo will ultimately lose after winning and losing some of the five games. However, people say that Lee Sedol earned USD 1 million price money for free, but I do not agree. I would have feared the result”.

Gu Li (9p): “Without any doubt, this has been an astonishing development, I believe it will defeat human in the future”. 

Shi Yue (9p): “It will probably be a good opportunity for us when the program reaches the level of a top player and is accessible by the general public. I would definitely play with it every day using different tactics in order to understand more about Go”. 

Changhyuk Yoo (9p, head coach of Korean Baduk Team): “If AlphaGo's current level is similar to the one it showed during the match with Fan Hui, Lee Sedol will easily beat AlphaGo. However, we are not sure how much progress AlphaGo has made during the six months after the Fan Hui match. I originally expected that it will take long to see AI catch up with humans in Go, but I was surprised to see AlphaGo win the match against Fan Hui”. 


The main and big unknown in all expectations worldwide about the outcome of the match is the relative strength of the current version of AlphaGo compared to the version that defeated Fan Hui. Was DeepMind able to further improve AlphaGo's way of playing structurally in the mean time to approach the caliber of a 9p? If not, predicting the outcome of the match seems to be no brainer.  

Based on the match with Fan Hui, several weaknesses in the way AlphaGo's plays have been suggested (e.g. Younggil An, 8p): a lack of understanding the concept of sente, no insight in the principle of aji, misjudging   complicated moves with delayed consequences further on in the game, problems with complicated and big ko's, and an clear absence of 'creativity' by just following common patterns over and over again: AlphaGo mimics the play of professionals and then follows usually standard patterns that may not turn out to be optimal in specific positions that demand for precise, proficient, and perceptive deviations.  


Myungwan Kim (9p) also commented about AlphaGo's apparent lack of whole-board awareness: "while professional players are more creative and will vary their play more based on subtle differences in other part of the board, AlphaGo makes '5-dan mistakes' while struggling with the whole-board interconnectedness". This may be due to specific structure of the underlying models of the program: the convolutional neural networks AlphaGo are based on, are typically local in nature, and therefore don’t build a coherent whole-board understanding. Therefore, we don’t know yet how AlphaGo will do when the fighting gets more extended and complex, or when the board is more fluid and multiple local positions are not unresolved yet.

According to AI games programmer RĂ©mi Coulom AlphaGo cannot propagate information at a distance more than 13 points due to it's underlying neural network architecture (1 layer of 5x5, 11 layers of 3x3 convolution). So when there is a fight on one side of the board, AlphaGo is unlikely to assess and interpret correctly local positions on the other side of the board. Therefore, the program might getting into serious trouble in positions with important, non-local fields of tension (for instance, during multple fights occuring simultaneously all over the board.  
DeepMind research scientist Thore Graepel explains: “Although we have programmed this machine to play, we have no idea what moves it will come up with.It moves are an emergent phenomenon from the training. We just create the data sets and the training algorithms. But the moves AlphaGo then comes up with are out of our hands – and much better than we, as Go players, could come up with. The program's rather autonomous nature”.
 

DeepMind chief executive Demis Hassabis adds: “AlphaGo played itself, different versions of itself, millions and millions of times and each time got incrementally slightly better as it learned from its mistakes”.  About two weeks before the match, Aja Huang (6d) commented: “We are still preparing hard for the match, AlphaGo is getting stronger and stronger”. Learning and improving from its own matchplay experience means deep learning AlphaGo is now even stronger than when it beat European champion Fan Hui last year. Asked whether there is a ceiling to AlphaGo's learning abilities, Hassabis answered: “If there is one, we haven’t found it yet”.

Hence an important question is in what areas AlphaGo has been improved further by the DeepMind team? What smart and effective upgrades of AlphaGo over the past five months can we expect? Thereby, we roughly can make four classes of improvements (see also: 'Let it Go' event,  Leo Dorst, UvA): data, algorithms and software, training and learning, hardware and computing time:

Data

-selection of stronger pro games (not only ≥ 6d from KGS, also collections of prof games). It has been shown that with small improvements in the accuracy of reproducing prof moves, immediately big leaps forward can be made in playing strength.
-extension of AlphaGo with comprehensive joseki-, shape- and/or complex pattern libraries
-comprehensive analysis and training on Lee Sedol games specifically, perhaps targeting weaker elements in the way Lee Sedol plays (if these exist at all)

Algorithms and Software

-AlphaGo has learned from mistakes it made during the match against Fan Hui
-improvement of AlphaGo's algorithms for e.g.
move selection and board evaluation
-prevent / circumvent specific problematic  situations (for instance: complex ko situations)
-improve and/or extend feature filters in order to be able to represent positions better and in more detail, these determine whether a (subpart of a) position during a game against Lee Sedol is sufficiently precise being recognized and accurately classified by AlphaGo
-improvement of the balance between on one side the neural networks for move selection and board evaluation and on the other side precise computation using Monte Carlo Tree Search
-extension of the number of network layers in order to recognize new, more specific go-features 
-incorporate new ideas and concepts to increase the performance of AlphaGo's play

Training and Learning

-fine tuning and extension of AlphaGo's neural network training sessions
-extension of the number of studied go-positions (>60 million) and/or games played (e.g. against itself, ≥ 1.3 million) to increase the accuracy of selecting and playing winning moves by AlphaGo
-improvements in learning the value of go-moves, for instance by more detailed and more accurate backpropagation of the final game result during the training sessions

Hardware and Computing Time:

-extension of the number of conventional (> 1202 CPUs) and graphical processors (>176 GPUs) the distributed version of AlphaGo can use simultaneously
-increasing the thinking time / computation time)
(this used to be 1 hour per person during the Fan Hui match, and apparently will be now 2 hour, which will be particularly beneficial for AlphaGo, especially towards the endgame


Which of these improvements will be applied and whether they will compensate for the weaknesses in play of the earlier version of AlphaGo? It is hard to tell what another five months 24/7 complimentary neural network training can do for AlphaGo's way of playing. According to William Sanzenin (quora.com): “I’d expect the March version of AlphaGo to be significantly stronger than the one that played last October. Its reading is going to be even better. Its assessment of the global position will be improved”. 

In case the professionals’ observations about AlphaGo’s play do reflect limitations inherent to its system structure and / or approach, as a matter of course,  this is fatal  against Lee Sedol. Also, it is possible that the program simply hasn't studied enough top-prof game positions. Herewith, the generally accepted idea is that deep learning models are as good as the data you feed them (see also: AlphaGo under a Magnifying Glass). So pro games as learning material undoubtedly will increase substantially AlphaGo's playing strength.

Concerning the important question what will be the hardware used for the distributed version of AlphaGo during the match against Lee Sedol, Demis Hassabis tweeted: “We are using roughly same amount of compute power as in Fan Hui match: distributing search over further machines has diminishing returns”. That means AlphaGo will be using 1,202 CPUs and 176 GPUs. However, various sources, including the New Economist, state that “the actual version of AlphaGo will be using 1,920 CPUs and 280 GPUs which is broadly similar to the computer power used in the match with Fan Hui”.

Apart from the extended training of AlphaGo by tera playing against itself, the possibly extended hardware, and the increased thinking time, nobody actually knows precisely what detailed improvements in AlphaGo's play have been achieved  over the past five months.

Another issue is whether AlphaGo can adapt in real-time (during and / or after a game) to the way Lee Sedol plays. Hassabis: “This is about teaching and learning. One game is not enough data to learn from –for a machine-- and training takes an awful lot of time: training a new version of AlphaGo takes about 4 – 6 weeks”.

On March 8th, the evening before the match, on nearly all polls worldwide show similar forecasts: about 75 - 85 % of all voters is convinced Lee Sedol will win the match against AlphaGo. This is also in line with the predictions, in advance of the match, of a price contest among dutch Go players concerning the expected outcome (106 participants, including the strongest amateur players in Europe, organised in cooperation with chess and go shop 'het Paard ' and an ICT company). 


To summarize: there are many reasons to believe that the actual version of AlphaGo has climbed at least a few dan grades: the expected playing strength of AlphaGo is ≥ 8p with a base of expected improvements and the fact that AlphaGo's playing strength already was ~5p at time of the Fan Hui match.

In advance of the match, Lee Sedol is very self-confident about winning the match against AlphaGo. When invited for the match Lee Sedol reacted: “This is the first time a computer has challenged a human pro to an even game and I am privileged to be the one to play it. Regardless of the result, it will be a meaningful event in baduk history. I heard Google DeepMind's AI is surprisingly strong and getting stronger, but I am confident that I can win, at least this time”. 

In an interview in Yonhap News, Lee Sedol told he is confident of beating AlphaGo by a score of 5-0, or at least with 4-1, and that he actually accepted the challenge after only 5 minutes of thinking. Also, he stated:  "Of course, there would have been many updates of AlphaGo in the last four or five months, but that isn’t enough time to challenge me". And a couple of weeks before the match, in an inverview with Sohn Suk-hee, Lee Sedol stated that: "even if I beat AlphaGo by 4-1,
the Google DeepMind team has the right to claim its --de facto-- victory and to celebrate my defeat, or even that of mankind".