Who is Zhengwei Wu? Created both optimal control tutorials! Optimal control: tutorials Fun fact: The only one in lab who could solve this geared cube puzzle How should the brain act? Electrical and Computer Engineering by Zhengwei Wu, Xaq Pitkow, Shreya Saxena 1 Neuroscience Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 2 Perception-Action Loop Perception-action Loop Bayesian modeling provides key tools for formalizing and understanding the perception-action loop. Tues-Weds covered models for time series, and integrating measurements over time. Optimal control: choose actions to achieve greatest expected value. Thurs-Fri we're using those tools to go beyond perception and include action! Ut utility world state measurements actions beliefs U s s t s t t–1 t+1 mt–1 at–1 b t–1 mt at bt m t+1 bt+1 time Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 3 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 4 Week 2 ⦁ Day 4 ⦁ Tutorials Markov Decision Process (MDP) Markov Decision Process (MDP) sequential decision-making, maximize rewards sequential decision-making, maximize rewards current world state is observable — future is unknown current world state is observable — future is unknown st at state action p(st+1 |st , at ) transitions R(s,a) rewards π(a|s) policy Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control } value V(s) st state at action p(st+1 |st , at ) transitions R(s,a) rewards π(a|s) policy } Rt–1 value V(s) world state actions 5 6 Partially Observable MDP (POMDP) Partially Observable MDP (POMDP) Week 2 ⦁ Day 4 ⦁ Tutorials Current world state is uncertain: partial observations Current uncertainty known — future uncertainty unknown p(st+1 |st , at ) U(s,a) π(a|s) Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 8 t+1 at at–1 Current uncertainty known — future uncertainty unknown state action transitions utility policy s t t–1 Current world state is uncertain: partial observations st at 7 s s Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials Rt } value V(s) Week 2 ⦁ Day 4 ⦁ Tutorials Partially Observable MDP (POMDP) Partially Observable MDP (POMDP) Current world state is uncertain: partial observations Current world state is uncertain: partial observations Current uncertainty known — future uncertainty unknown Current uncertainty known — future uncertainty unknown bt =p(st |m0:t , a0:t ) belief state at action p(st+1 |st , at ) transitions U(s,a) utility π(a|s) policy } value Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control bt =p(st |m0:t , a0:t ) belief state mt measurement at action p(st+1 |st , at ) transitions U(s,a) utility π(a|s) policy V(s) Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 9 10 Partially Observable MDP (POMDP) Partially Observable MDP (POMDP) V(s) Week 2 ⦁ Day 4 ⦁ Tutorials Current world state is uncertain: partial observations Current world state is uncertain: partial observations Current uncertainty known — future uncertainty unknown Current uncertainty known — future uncertainty unknown bt =p(st |m0:t , a0:t ) belief state mt measurement at action p(bt+1 |bt , at ) belief transitions U(s,a) utility π(a|s) policy Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 11 } value } value bt =p(st |m0:t , a0:t ) belief state mt measurement at action p(bt+1 |bt , at ) belief transitions U(b,a) expected utility π(a|s) policy V(s) Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 12 } value V(s) Week 2 ⦁ Day 4 ⦁ Tutorials Partially Observable MDP (POMDP) Partially Observable MDP (POMDP) Current world state is uncertain: partial observations Current world state is uncertain: partial observations Current uncertainty known — future uncertainty unknown Current uncertainty known — future uncertainty unknown bt =p(st |m0:t , a0:t ) belief state mt measurement at action p(bt+1 |bt , at ) belief transitions U(b,a) expected utility π(a|b) policy } value Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control bt =p(st |m0:t , a0:t ) belief state mt measurement at action p(bt+1 |bt , at ) belief transitions U(b,a) expected utility π(a|b) policy V(s) Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 13 V(b) Week 2 ⦁ Day 4 ⦁ Tutorials 14 MDPs are interactive Markov Models POMDPs are interactive HMMs Partially Observable MDP (POMDP) Current world state is uncertain: partial observations Current uncertainty known — future uncertainty unknown bt =p(st |m0:t , a0:t ) mt belief state measurement at action p(bt+1 |b t, at) belief transitions U(b,a) expected utility π(a|b) policy } Ut utility world state value V(b) measurements actions beliefs Weds you found posterior of current state using measurements: • Telegraph process (binary HMM) • Continuous linear dynamical system (Kalman filter) U s t st s t–1 t+1 mt–1 at–1 b mt at bt m Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Today you'll take actions based on those HMM inferences t+1 b t–1 15 } value t+1 Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 16 Week 2 ⦁ Day 4 ⦁ Tutorials Optimal Control: Open and Closed Loop Overview of tutorials for Optimal Control day Open loop control solves POMDP without measurements 1: Fishing • Control a Binary latent variable • Continues from Binary HMM • Connects to Reinforcement learning (tomorrow) Closed loop control solves POMDP with measurements 2: Flying • Linear optimal control of Continuous state • Continues from Linear Dynamics, and Kalman Filter Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 17 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 18 Tutorial 1: Gone Fishin' Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 19 Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei 20 ⦁ Day 4 ⦁ Tutorials Zhengwei ⦁ Day 4 ⦁ Tutorials 21 Zhengwei ⦁ Day 4 ⦁ Tutorials 22 basic Bayes s world state t mt measurements Markov Model s t –1 s t s t+1 Hidden Markov Model s s t t –1 m t –1 s t+1 mt mt 1 Markov Decision Process + world state actions measurements Zhengwei 23 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 24 s s s t t –1 m t –1 a t –1 m t a t+1 t Week 2 ⦁ Day 4 ⦁ Tutorials m t +1 basic Bayes world state measurements basic Bayes s s world state t mt t measurements mt Hidden Markov Model s s t t–1 s t+1 m mt t–1 mt+1 Markov Decision Process world state actions measurements Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control s s s t t–1 m t–1 a t–1 m t a t+1 m t Week 2 ⦁ Day 4 ⦁ Tutorials 25 t +1 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 26 basic Bayes world state measurements s t mt State space s position of fish State space s Measurements m position of fish Fishing on Right side... Measurements m Fishing on Right side... s = Left m = Catch fish m = Nope s= Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 27 Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 28 s = Left 0.1 0.9 s = Right 0.5 0.5 Week 2 ⦁ Day 4 ⦁ Tutorials basic Bayes world state measurements s t mt Hidden Markov Model Hidden Markov Model s st–1 s t s t–1 ... t+1 m t–1 mt m mt+1 t–11 Markov Decision Process world state actions measurements s s m t–1 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control a t–1 t mt s ... t+1 mt+11 s t t–1 ss m t a t+1 m t Week 2 ⦁ Day 4 ⦁ Tutorials 29 t +1 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 30 ... m t–1 s mt Special case: Fixed latent ... s ... m t–1 mt+1 mt ... mt+1 Hidden Markov Model s s t t–1 m t–1 s t+1 mt Changing latent State space s position of fish mt+1 asurements m ng on Right side... Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 31 Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 32 Enough! I know already Week 2 ⦁ Day 4 ⦁ Tutorials Tutorial 1: Gone Fishin' Zhengwei ⦁ Day 4 ⦁ Tutorials 33 Zhengwei 35 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 34 ⦁ Day 4 ⦁ Tutorials Zhengwei 36 ⦁ Day 4 ⦁ Tutorials Zhengwei ⦁ Day 4 ⦁ Tutorials 37 Zhengwei 38 How should you act? How should you act? Optimal control problem: where should I go to catch the most fish? Optimal control problem: where should I go to catch the most fish? Must define: • states s • actions a • dynamics D • Utility U(s,a) Maximize total expected Utility = Reward – Cost Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 39 ⦁ Day 4 ⦁ Tutorials Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 40 Week 2 ⦁ Day 4 ⦁ Tutorials Actions State space Two state variables: Left You control your own position. You can stay or switch sides. Right position of fish Left Switching sides costs energy. Right position of you Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 41 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 42 Actions Actions You control your own position. You can stay or switch sides. You control your own position. You can stay or switch sides. Switching sides costs energy. Left Switching sides costs energy. Right Let's measure that in units of fish. position of you Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 43 Week 2 ⦁ Day 4 ⦁ Tutorials Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 44 Left Right Whew! That was 3-fish tiring. position of you Week 2 ⦁ Day 4 ⦁ Tutorials Your turn: look at the dynamics in the code! Dynamics: Telegraph process p( Fish are on Left | Fish were on Left ) Plot the telegraph process for different p(stay). 1 fish p(st+1 = Left|s t = Left) = 0.95 p(st+1 = Right|st = Left) = 0.05 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 0 Stay Switch Week 2 ⦁ Day 4 ⦁ Tutorials 45 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 46 Dynamics: Telegraph process Dynamics: Telegraph process P( fish there at time t | fish there at time 0 ) Right 1 time prob(stay) timescale 0.9 10 0.95 20 0.98 50 p 1 / (1–p) 47 .9 .95 .98 0 Left Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 1/2 Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 48 time t Week 2 ⦁ Day 4 ⦁ Tutorials Measurements = Rewards Your turn: catch some fish! P( catch fish | you are on the same side as the fish ) Plot the sequence of caught fish when Caught fish? YES NO p( catch fish | on other side ) p( catch fish | on same side ) P( catch fish | you are on a different side from the fish ) Caught fish? YES NO Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 49 Week 2 ⦁ Day 4 ⦁ Tutorials 50 Measurements = Rewards Where are the fish? P( catch fish | you are on the same side as the fish ) You only learn indirectly from what you've caught: • This is an HMM! • Binary latent state • Binary observations (unlike DDM) • Only observe on one side Caught fish? YES time NO Well, I caught one on this side three hours ago... P( catch fish | you are on a different side from the fish ) Caught fish? YES Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 51 time NO Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 52 Week 2 ⦁ Day 4 ⦁ Tutorials Where are the fish? Where are the fish? You only learn indirectly from what you've caught: • This is an HMM! • Binary latent state • Binary observations (unlike DDM) • Only observe on one side You only learn indirectly from what you've caught: • This is an HMM! • Binary latent state • Binary observations (unlike DDM) • Only observe on one side Well, I caught one on this side three hours ago... Recall HMM posterior equations: Recall HMM posterior equations: 1 p(st |m1:t , a1:t−1 ) = Z p(mt |st ) ∑ p(st |st−1 , at−1 )p(st−1 |m1:t−1 , a1:t−2 ) 1 p(st |m1:t ) = Z p(mt |st ) ∑ p(st |st−1 )p(st−1 |m1:t−1 ) s st−1 t−1 This posterior quantifies our "belief" about where the fish are. This posterior quantifies our "belief" about where the fish are. Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Well, I caught one on this side three hours ago... Week 2 ⦁ Day 4 ⦁ Tutorials 53 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 54 Where are the fish? Your turn: Plot belief on one side only Plot belief about fish location for telegraph process Compute and plot belief • p( fish state = Right | history of measurements on Right ) Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 55 Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 56 Week 2 ⦁ Day 4 ⦁ Tutorials How should you act? How should you act? World State Utility: • Utility of fish = 1 • Utility of state = Utility of fish x Probability of catching fish in state Belief State Utility (assuming accurate beliefs): • Utility of fish = 1 • Utility of state = Utility of fish x Probability of catching fish in state x Belief that fish are on your side + Utility of other side x Belief that fish are on other side Action Utility: • Utility of staying = 0 • Utility of switching = –Travel cost Action Utility: • Utility of staying = 0 • Utility of switching = –Travel cost Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 57 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 58 Policy: what to do in any situation Policy: what to do in any situation Here, the 'situation' is the current State: • Your position styou • Your belief about the fish position, btfish = p( stfish = Right | m1:t , a1:t ) Here, the 'situation' is the current State: • Your position styou • Your belief about the fish position, btfish = p( stfish = Right | m1:t , a1:t ) There are only two actions: stay or switch. There are only two actions: stay or switch. Policy: switch when chance that you're on the same side as the fish is too low. How low? Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 59 Week 2 ⦁ Day 4 ⦁ Tutorials Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 60 at stay switch you're on θ wrong side bt you're on same side Week 2 ⦁ Day 4 ⦁ Tutorials Evaluate policy Your turn: Implement threshold policy Complete the code for the policy Value is the total expected utility 1 you fish Value = T ∑ U(st , bt , at ) t = fraction of time on each side x fraction of time rewarded – switching rate x cost per switch Then run the interactive demo that uses your new policy Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 61 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 62 Your turn: Evaluate policy Your turn: Evaluate policy Compute Value = total expected utility, for one threshold. Compute Value = total expected utility, for different thresholds. Plot it for all thresholds in a range using: Then pick the optimal threshold! Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 63 Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 64 Week 2 ⦁ Day 4 ⦁ Tutorials Sensitivity of optimal policy Your turn: Evaluate policy Compute Value = total expected utility, for different thresholds. How does the optimal control change with: • travel cost? • fish dynamics? • odds of catching fish? Plot it for all thresholds in a range. Then pick the optimal threshold! Value ? threshold Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control Week 2 ⦁ Day 4 ⦁ Tutorials 65 Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 66 From discrete to continuous control Tutorial 2 — Continuous control Tutorial 1 covered control in a POMDP with: • discrete states • discrete actions • discrete measurements Again control solves a POMDP, now combining: • Control principles from Tutorial 1 • Kalman filtering from W2D3. We will gradually build up to full Linear Quadratic Gaussian control: • Start with open-loop control • Continue with closed-loop but full observability • Finally use partial observability and closed-loop control Tutorial 2 will cover control in a POMDP with: • continuous (gaussian) states • continuous actions • continuous measurements Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 67 Week 2 ⦁ Day 4 ⦁ Tutorials Week 2 ⦁ Day 4 ⦁ Tutorials Zhengwei Wu, Xaq Pitkow, Shreya Saxena ⦁ Optimal Control 68 Week 2 ⦁ Day 4 ⦁ Tutorials
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )