Correction: Exploration-Exploitation Guided Drone Dispatch Strategies in Uncertain Environments

Sihang Wei, Soham Phanse, Teja Koduru, Minghao Chen, Derrick W. Yeo, Max Z. Li · 2024

Modified: In this work, we propose an exploration-exploitation framework that leverages differing fleet priorities to simultaneously explore (using lower-priority missions) and exploit (used by using higher-priority missions) the unknown environment.Using the example application of urban wind flow fields, we show that a Thompson sampling-inspired exploration approach produces more optimal mission results, as measured by regret-like performance metrics. III-D. Exploration-Exploitation with Greedy Sampling-Optimal Dispatch Policy• Modified: Simultaneously, we also integrate an approach to exploit the obtained information to make data-driven decisions to improve fleet efficiency and reduce in-flight risks and costs.• Modified: When a sample is drawn at a waypoint 𝑤 𝑖 , a random sample is drawn from the true underlying wind distribution (denoted by GT 𝑖 from now on) referenced from the data set as described in Section IV.B.We denote this true underlying wind velocity distribution by GT .Hence, GT 𝑖 at waypoint 𝑖 is comprised of GT 𝑖,𝑥 and GT 𝑖,𝑦 , which denote the true wind velocity distribution in the 𝑥 and 𝑦 directions, respectively.Both of these components are normally distributed with true wind field means and variances denoted by 𝜇 𝑡 ,𝑖,𝑥 , 𝜇 𝑡 ,𝑖,𝑦 , 𝜎 2 𝑡 ,𝑖,𝑥 , and 𝜎 2 𝑡 ,𝑖,𝑦 .• Modified: We term this estimate as the current fleet belief and denote it by GP.Hence, GP 𝑖 denotes the current fleet belief at waypoint 𝑖, and is comprised of GP 𝑖,𝑥 and GP 𝑖,𝑦 , which denote the estimated wind velocity distributions in the 𝑥 and 𝑦 directions, respectively, at waypoint 𝑖.These components are normally distributed with the estimated wind field means and variances denoted by 𝜇 𝑐,𝑖,𝑥 , 𝜇 𝑐,𝑖,𝑦 , 𝜎 2 𝑐,𝑖,𝑥 , and 𝜎 2 𝑐,𝑖,𝑦 at waypoint 𝑖.It signifies that based on the samples collected so far, the fleet believes that the underlying distribution at waypoint 𝑤 𝑖 is given by GP 𝑖 ∼ N( 𝝁 𝑐,𝑖 , 𝚺 𝑐,𝑖 ).• In Reward step, modified: Note that 𝜇 𝑐,𝑖,𝑥 , 𝜇 𝑐,𝑖,𝑦 , and 𝜎 𝑐,𝑖,𝑥 , 𝜎 𝑐,𝑖,𝑦 are the current belief's wind velocity means and variances at waypoint 𝑖, in the 𝑥 and 𝑦 directions, respectively.We will see how the prior fleet belief can be used to estimate the costs incurred at a particular waypoint in later sections and how these costs, in turn, can be used to make decisions regarding which drone to send next in the dispatch queue.• In the initialization step: After defining the prior as above, we further collect 𝜆 samples at each waypoint and compute sample means and sample sum of squares.The prior and the sample parameters will be used to compute the posterior distributions (described in later sections) of the means and variances of the current belief, at each waypoint, in both 𝑥 and 𝑦 directions.The initialization process proceeds formulated as follows: • Initialization step (e):1

Read the paper · More papers on PaperTik