{"id":98500,"date":"2020-10-27T13:49:17","date_gmt":"2020-10-27T10:49:17","guid":{"rendered":"https:\/\/en.buradabiliyorum.com\/what-the-hell-is-reinforcement-learning-and-how-does-it-work\/"},"modified":"2020-10-27T13:49:17","modified_gmt":"2020-10-27T10:49:17","slug":"what-the-hell-is-reinforcement-learning-and-how-does-it-work","status":"publish","type":"post","link":"https:\/\/buradabiliyorum.com\/en\/what-the-hell-is-reinforcement-learning-and-how-does-it-work\/","title":{"rendered":"#What the hell is reinforcement learning and how does it work?"},"content":{"rendered":"<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_88 counter-hierarchy ez-toc-counter ez-toc-custom ez-toc-container-direction\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<label for=\"ez-toc-cssicon-toggle-item-6ac4375a4b943\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #dd3333;color:#dd3333\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #dd3333;color:#dd3333\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6ac4375a4b943\" checked aria-label=\"Toggle\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/buradabiliyorum.com\/en\/what-the-hell-is-reinforcement-learning-and-how-does-it-work\/#Reinforcement_learning_challenges\" >Reinforcement learning challenges<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/buradabiliyorum.com\/en\/what-the-hell-is-reinforcement-learning-and-how-does-it-work\/#Applications_areas_of_reinforcement_learning\" >Applications areas of reinforcement learning<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/buradabiliyorum.com\/en\/what-the-hell-is-reinforcement-learning-and-how-does-it-work\/#Conclusion_When_should_you_use_RL\" >Conclusion: When should you use RL?<\/a><\/li><\/ul><\/nav><\/div>\n<p>&#8220;<strong>#What the hell is reinforcement learning and how does it work?<\/strong>&#8221;<br \/>\n<img decoding=\"async\" src=\"https:\/\/cdn0.tnwcdn.com\/wp-content\/blogs.dir\/1\/files\/2020\/10\/1-19-796x417.jpg\" \/><\/p>\n<div>\n                                Reinforcement learning is a subset of machine learning. It enables an agent to learn through the consequences of actions in a specific environment. It can be used to teach a robot new tricks, for example.<\/p>\n<p>Reinforcement learning is a behavioral learning model where the algorithm provides data analysis feedback, directing the user to the best result.<\/p>\n<p>It differs from other forms of supervised learning because the sample data set does not train the machine. Instead, it learns by trial and error. Therefore, a <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/watch-movies-tv-seriess\/\" data-internallinksmanager029f6b8e52c=\"8\" title=\"Watch Movies &amp; TV Series\" target=\"_blank\" rel=\"noopener\">series<\/a> of right decisions would strengthen the method as it better solves the problem.<\/p>\n<p>Reinforced learning is similar to what we humans have when we are children. We all went through the learning reinforcement \u2014 when you started crawling and tried to get up, you fell over and over, but your parents were there to lift you and teach you.<\/p>\n<p>It is teaching based on experience, in which the machine must deal with what went wrong before and look for the right <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/download-scripts-themes-apps\/\" data-internallinksmanager029f6b8e52c=\"9\" title=\"Download Scripts &amp; Themes &amp; Apps\" target=\"_blank\" rel=\"noopener\">app<\/a>roach.<\/p>\n<p>Although we don\u2019t describe the reward policy \u2014 that is, the <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/game\/\" data-internallinksmanager029f6b8e52c=\"7\" title=\"Game\" target=\"_blank\" rel=\"noopener\">game<\/a> rules \u2014 we don\u2019t give the model any tips or advice on how to solve the game. It is up to the model to figure out how to execute the task to optimize the reward, beginning with random testing and sophisticated tactics.<\/p>\n<p>By exploiting research power and multiple attempts, reinforcement learning is the most successful way to indicate computer imagination. Unlike humans, artificial intelligence will gain knowledge from thousands of side games. At the same time, a reinforcement learning algorithm runs on robust computer infrastructure.<\/p>\n<p>An example of reinforced learning is the recommendation on <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/social-mediaa\/\" data-internallinksmanager029f6b8e52c=\"1\" title=\"Social Media\" target=\"_blank\" rel=\"noopener\">Youtube<\/a>, for example. After watching a video, the platform will show you similar titles that you believe you will like. However, suppose you start watching the recommendation and do not finish it. In that case, the machine understands that the recommendation would not be a good one and will try another approach next time.<\/p>\n<p><em>[Read: <span class=\"c-message_attachment__title\"><span dir=\"auto\">What audience intelligence data tells us about the 2020 US presidential election<\/span>]<\/span><\/em><\/p>\n<h2><span class=\"ez-toc-section\" id=\"Reinforcement_learning_challenges\"><\/span>Reinforcement learning challenges<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Reinforcement learning\u2019s key challenge is to plan the simulation environment, which relies heavily on the task to be performed. When trained in Chess, Go, or Atari games, the simulation environment preparation is relatively easy. Building a model capable of driving an autonomous car is key to creating a realistic prototype before letting the car ride the street. The model must decide how to break or prevent a collision in a safe environment. Transferring the model from the training setting to the real world becomes problematic.<\/p>\n<p>Scaling and modifying the agent\u2019s neural network is another problem. There is no way to connect with the network except by incentives and penalties. This may lead to disastrous forgetfulness, where gaining new information causes some of the old knowledge to be removed from the network. In other words, we must keep learning in the agent\u2019s \u201cmemory.\u201d<\/p>\n<p>Another difficulty is reaching a great location \u2014 that is, the agent executes the mission as it is, but not in the ideal or required manner. A \u201chopper\u201d jumping like a kangaroo instead of doing what is expected of him is a perfect example. Finally, some agents can maximize the prize without completing their mission.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Applications_areas_of_reinforcement_learning\"><\/span>Applications areas of reinforcement learning<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p><strong>Games<\/strong><\/p>\n<p>RL is so well known today because it is the conventional algorithm used to solve different games and sometimes achieve superhuman performance.<\/p>\n<p>The most famous must be AlphaGo and AlphaGo Zero. AlphaGo, trained with countless human games, has achieved superhuman performance using the Monte Carlo tree value research and value network (MCTS) in its policy network. However, the researchers tried a purer approach to RL \u2014 training it from scratch. The researchers left the new agent, AlphaGo Zero, to play alone and finally defeat AlphaGo 100\u20130.<\/p>\n<p><strong>Personalized recommendations<\/strong><\/p>\n<p>The work of <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/news\/\" data-internallinksmanager029f6b8e52c=\"2\" title=\"News\" target=\"_blank\" rel=\"noopener\">news<\/a> recommendations has always faced several challenges, including the dynamics of rapidly changing news, users who tire easily, and the Click Rate that cannot reflect the user retention rate. Guanjie et al. applied RL to the news recommendation system in a document entitled \u201cDRN: A Deep Reinforcement Learning Framework for News Recommendation\u201d to tackle problems.<\/p>\n<p>In practice, they built four categories of resources, namely: A) user resources, B) context resources such as environment state resources, C) user news resources, and D) news resources such as action resources. The four resources were inserted into the Deep Q-Network (DQN) to calculate the Q value. A news list was chosen to recommend based on the Q value, and the user\u2019s click on the news was part of the reward the RL agent received.<\/p>\n<p>The authors also employed other techniques to solve other challenging problems, including memory repetition, survival models, Dueling Bandit Gradient Descent, and so on.<\/p>\n<p><strong>Resource management in computer clusters<\/strong><\/p>\n<p>Designing algorithms to allocate limited resources to different tasks is challenging and requires human-generated heuristics.<\/p>\n<p>The article \u201cResource management with deep reinforcement learning\u201d explains how to use RL to automatically learn how to allocate and schedule computer resources for jobs on hold to minimize the average job (task) slowdown.<\/p>\n<p>The state-space was formulated as the current resource allocation and the resource profile of jobs. For the action space, they used a trick to allow the agent to choose more than one action at each stage of time. The reward was the sum of (-1 \/ job duration) across all jobs in the system. Then they combined the REINFORCE algorithm and the baseline value to calculate the policy gradients and find the best policy parameters that provide the probability distribution of the actions to minimize the objective.<\/p>\n<p><strong>Traffic light control<\/strong><\/p>\n<p>In the article \u201cMulti-agent system based on reinforcement learning to control network traffic signals,\u201d the researchers tried to design a traffic light controller to solve the congestion problem. Tested only in a simulated environment, their methods showed results superior to traditional methods and shed light on multi-agent RL\u2019s possible uses in traffic systems design.<\/p>\n<p>Five agents were placed in the five intersections traffic network, with an RL agent at the central intersection to control traffic signaling. The state was defined as an eight-dimensional vector, with each element representing the relative traffic flow of each lane. Eight options were available to the agent, each representing a combination of phases, and the reward function was defined as a reduction in delay compared to the previous step. The authors used DQN to learn the Q value of {state, action} pairs.<\/p>\n<p><strong>Robotics<\/strong><\/p>\n<p>There is an incredible job in the application of RL in robotics. We recommend reading this paper with the result of RL research in robotics. In this other work, the researchers trained a robot to learn policies to map raw video images to the robot\u2019s actions. The RGB images were fed into a CNN, and the outputs were the engine torques. The RL component was policy research guided to generate training data from its state distribution.<\/p>\n<p><strong>Web systems configuration<\/strong><\/p>\n<p>There are more than 100 configurable parameters in a Web System, and the process of adjusting the parameters requires a qualified operator and several tracking and error tests.<\/p>\n<p>The article \u201cA learning approach by reinforcing the self-configuration of the online Web system\u201d showed the first attempt in the domain on how to autonomously reconfigure parameters in multi-layered web systems in dynamic VM-based environments.<\/p>\n<p>The reconfiguration process can be formulated as a finite MDP. The state-space was the system configuration; the action space was {increase, decrease, maintain} for each parameter. The reward was defined as the difference between the intended response time and the measured response time. The authors used the Q-learning algorithm to perform the task.<\/p>\n<p>Although the authors used some other technique, such as policy initialization, to remedy the large state space and the computational complexity of the problem, instead of the potential combinations of RL and neural network, it is believed that the pioneering work prepared the way for future research in this area\u2026<\/p>\n<p><strong>Chemistry<\/strong><\/p>\n<p>RL can also be applied to optimize chemical reactions. Researchers have shown that their model has outdone a state-of-the-art algorithm and <a href=\"https:\/\/buradabiliyorum.com\/en\/category\/general\/\" data-internallinksmanager029f6b8e52c=\"3\" title=\"General\" target=\"_blank\" rel=\"noopener\">general<\/a>ized to different underlying mechanisms in the article \u201cOptimizing chemical reactions with deep reinforcement learning.\u201d<\/p>\n<p>Combined with LSTM to model the policy function, agent RL optimized the chemical reaction with the Markov decision process (MDP) characterized by {S, A, P, R}, where S was the set of experimental conditions ( such as temperature, pH, etc.), A was the set of all possible actions that can change the experimental conditions, P was the probability of transition from the current condition of the experiment to the next condition and R was the reward that is a function of the state.<\/p>\n<p>The application is excellent for demonstrating how RL can reduce time and trial and error work in a relatively stable environment.<\/p>\n<p><strong>Auctions and advertising<\/strong><\/p>\n<p>Researchers at Alibaba Group published the article \u201c<a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"de fg\" href=\"https:\/\/arxiv.org\/pdf\/1802.09756.pdf\">Real-time auctions with multi-agent reinforcement learning in display advertising<\/a>.\u201d They stated that their cluster-based distributed multi-agent solution (DCMAB) has achieved promising results and, therefore, plans to test the Taobao platform\u2019s life.<\/p>\n<p>Generally speaking, the Taobao ad platform is a place for marketers to bid to show ads to customers. This can be a problem for many agents because traders bid against each other, and their actions are interrelated. In the article, merchants and customers were grouped into different groups to reduce computational complexity. The agents\u2019 state-space indicated the agents\u2019 cost-revenue status, the action space was the (continuous) bid, and the reward was the customer cluster\u2019s revenue.<\/p>\n<p><strong>Deep learning<\/strong><\/p>\n<p>More and more attempts to combine RL and other deep learning architectures can be seen recently and have shown impressive results.<\/p>\n<p>One of RL\u2019s most influential jobs is Deepmind\u2019s pioneering work to combine CNN with RL. In doing so, the agent can \u201csee\u201d the environment through high-dimensional sensors and then learn to interact with it.<\/p>\n<p><a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"de fg\" href=\"https:\/\/www.cs.toronto.edu\/~vmnih\/docs\/dqn.pdf\">CNN with RL<\/a> are other combinations used by people to try new ideas. RNN is a type of neural network that has \u201cmemories.\u201d When combined with RL, RNN offers agents the ability to memorize things. For example, they combined LSTM with RL to create a deep recurring Q network (DRQN) for playing Atari 2600 games. They also used\u00a0<a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"de fg\" href=\"https:\/\/arxiv.org\/pdf\/1507.06527.pdf\">LSTM with RL<\/a> to solve problems in optimizing chemical reactions.<\/p>\n<p>Deepmind showed how to use <a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"de fg\" href=\"https:\/\/arxiv.org\/pdf\/1804.01118.pdf\">generative models and RL<\/a> to generate programs. In the model, the adversely trained agent used the signal as a reward for improving actions, rather than propagating gradients to the entry space as in GAN training. Incredible, isn\u2019t it?<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Conclusion_When_should_you_use_RL\"><\/span>Conclusion: When should you use RL?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Reinforcement is done with rewards according to the decisions made; it is possible to learn continuously from interactions with the environment at all times. With each correct action, we will have positive rewards and penalties for incorrect decisions. In the industry, this type of learning can help optimize processes, simulations, monitoring, maintenance, and the control of autonomous systems.<\/p>\n<p id=\"3e64\" class=\"ii ij fj ik b gi ja il im gl jb in io ip jc iq ir is jd it iu iv je iw ix iz dy gg\" data-selectable-paragraph=\"\"><a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"de fg\" href=\"https:\/\/bons.ai\/blog\/ai-reinforcement-learning-strategy-industrial-systems\">Some criteria<\/a> can be used in deciding where to use reinforcement learning:<\/p>\n<ul class=\"\">\n<li id=\"88bf\" class=\"ii ij fj ik b gi ja il im gl jb in io ip jc iq ir is jd it iu iv je iw ix iz kp kq kr gg\" data-selectable-paragraph=\"\">When you want to do some simulations given the complexity, or even the level of danger, of a given process.<\/li>\n<li id=\"efd2\" class=\"ii ij fj ik b gi ks il im gl kt in io ip ku iq ir is kv it iu iv kw iw ix iz kp kq kr gg\" data-selectable-paragraph=\"\">To increase the number of human analysts and domain experts on a given problem. This type of approach can <a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"de fg\" href=\"https:\/\/hackernoon.com\/reinforcement-learning-and-supervised-learning-a-brief-comparison-1b6d68c45ffa\">imitate human reasoning instead of learning the best possible strategy<\/a>.<\/li>\n<li id=\"6cae\" class=\"ii ij fj ik b gi ks il im gl kt in io ip ku iq ir is kv it iu iv kw iw ix iz kp kq kr gg\" data-selectable-paragraph=\"\">When you have a good reward definition for the learning algorithm, you can calibrate correctly with each interaction so that you have more positive than negative rewards.<\/li>\n<li id=\"edce\" class=\"ii ij fj ik b gi ks il im gl kt in io ip ku iq ir is kv it iu iv kw iw ix iz kp kq kr gg\" data-selectable-paragraph=\"\">When you have <a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"de fg\" href=\"https:\/\/hackernoon.com\/reinforcement-learning-and-supervised-learning-a-brief-comparison-1b6d68c45ffa\">little data<\/a> about a particular problem.<\/li>\n<\/ul>\n<p>In addition to industry, reinforcement learning is used <a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"de fg\" href=\"https:\/\/www.oreilly.com\/ideas\/practical-applications-of-reinforcement-learning-in-industry\">in various fields such<\/a> as education, health, finance, image, and text recognition.<\/p>\n<hr\/>\n<p><em>This article was written by <\/em>Jair Ribeiro<span class=\"hp\"> <\/span><em>and was originally published on <a rel=\"nofollow noopener noreferrer\" target=\"_blank\" href=\"https:\/\/towardsdatascience.com\/\">Towards Data Science<\/a>. You can read it <a rel=\"nofollow noopener noreferrer\" target=\"_blank\" href=\"https:\/\/towardsdatascience.com\/about-reinforcement-learning-2ff0dafe9b75\">here<\/a>.\u00a0<\/em><\/p>\n<p class=\"c-post-pubDate\">\n                                    Published October 27, 2020 \u2014 10:49 UTC\n                                <\/p>\n<\/p><\/div>\n<p><script data-src=\"https:\/\/connect.facebook.net\/en_US\/sdk.js#xfbml=1&amp;appId=378011798897423&amp;version=v2.6\" id=\"socialSrcFacebook\" type=\"text\/template\"><\/script><\/p>\n<blockquote>\n<p style=\"text-align: center;\">For forums sites go to <span style=\"color: #ff9900;\"><a style=\"color: #ff9900;\" href=\"https:\/\/forum.buradabiliyorum.com\/\" target=\"_blank\" rel=\"noopener noreferrer\">Forum.BuradaBiliyorum.Com<\/a><\/span><\/strong><\/p>\n<\/blockquote>\n<blockquote>\n<p style=\"text-align: center;\"><strong>If you want to read more like this article, you can visit our <span style=\"color: #ff9900;\"><a style=\"color: #ff9900;\" href=\"https:\/\/en.buradabiliyorum.com\/technology\/\" target=\"_blank\" rel=\"noopener noreferrer\">Technology category.<\/a><\/span><\/strong><\/p>\n<\/blockquote>\n<p><span style=\"color: black;\"><a style=\"color: #ff9900;\" href=\"https:\/\/thenextweb.com\/neural\/2020\/10\/27\/what-the-hell-is-reinforcement-learning-and-how-does-it-work-syndication\/\" target=\"_blank\" rel=\"noopener noreferrer\">Source<\/a><\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>&#8220;#What the hell is reinforcement learning and how does it work?&#8221; Reinforcement learning is a subset of machine learning. It enables an agent to learn through the consequences of actions in a specific environment. It can be used to teach a robot new tricks, for example. Reinforcement learning is a behavioral learning model where the&#8230;<\/p>\n","protected":false},"author":1,"featured_media":98501,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/img-cdn.tnwcdn.com\/image\/neural?filter_last=1&fit=1280,640&url=https:\/\/cdn0.tnwcdn.com\/wp-content\/blogs.dir\/1\/files\/2020\/10\/1-19.jpg&signature=2f0aa931a6b6a3da1e5e5f67b6746c35","fifu_image_alt":"","footnotes":""},"categories":[18],"tags":[],"class_list":["post-98500","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology"],"_links":{"self":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts\/98500","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/comments?post=98500"}],"version-history":[{"count":0,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/posts\/98500\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/media\/98501"}],"wp:attachment":[{"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/media?parent=98500"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/categories?post=98500"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/buradabiliyorum.com\/en\/wp-json\/wp\/v2\/tags?post=98500"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}