Warning: Trying to access array offset on false in /var/www/hswl/user4/html/IncludeFiles/news-detail-common.php on line 350
[Shunli Smart Logistics Seminar] How to Solve the Challenges of AGV and Cargo Elevator Scheduling in Multi-Floor Workshops? Deep Reinforcement Learnin
Guangdong Shunli Intelligent Logistics Equipment Co., Ltd.
CN
1
Current Location: Home > News Center > Industry News > [Shunli Smart Logistics Seminar] How to Solve the Challenges of AGV and Cargo Elevator Scheduling in Multi-Floor Workshops? Deep Reinforcement Learnin

Hot KeywordsKeywords

Contact UsContact Us

Guangdong Shunli Intelligent Logistics Equipment Co., Ltd.

Mobile:13631707130

Email:marketing@dgsunli.com

Tel:400-090-1058

Website:www.gdsunli.com

Address:No. 115, Gaojiu Section, Beiwang Road, Gaobu Town, Dongguan City, Guangdong Province

[Shunli Smart Logistics Seminar] How to Solve the Challenges of AGV and Cargo Elevator Scheduling in Multi-Floor Workshops? Deep Reinforcement Learnin

2026-08-13 17:27:27
6 Views

[This Week's Topic]

Resolution of the bottleneck in the logistics scheduling of the fifth-floor workshop:

How to make AGVs, cargo elevators and production rhythms work in harmony?

This issue's technical director: Li Weijian


I. Abstract

This is a set of Shunli Intelligent deep reinforcement learning logistics collaborative scheduling algorithm tailored for the five-story automotive component production scenario: It not only assigns routes to AGVs, but also takes into account processing cycle time, material completeness, work-in-process inventory, vehicle load, and elevator status, continuously determining "who will transport, what to transport, where to deliver, how much to transport at one time, and when to cross floors", to facilitate the smooth production of more automotive component assemblies within a unit time.


Traditional focus

The focus of this algorithm

Does a single AGV travel a shorter distance?

Has the entire production chain advanced more smoothly towards the final assembly?

Is the freight elevator operating less frequently?

Has the freight elevator been utilized in a timely manner for truly valuable cross-floor tasks?

Does a certain workstation maintain a high inventory?

Are all the upstream and downstream components complete? Has the situation been avoided where "there is overstocking at the upstream stage and shortage at the downstream stage"?

Complete one transportation task at a time

When process compatibility and capacity permit, combine and transport multiple small-batch orders together

This article uses a five-layer production scenario of automotive reducer components, which is constructed for technical illustration. This scenario does not correspond to the actual production line of any specific enterprise. Its purpose is to explain how the algorithm handles the collaborative scheduling problem of multi-floor, multi-station, multi-material and shared elevator through the common processing procedures of automotive gears and shaft parts.

II. Content Overview

1. The fifth-floor automotive parts workshop: Why logistics scheduling becomes a production bottleneck

2. Optimization objective: Not to make the equipment busier, but to ensure that more assemblies can be released from the production line in a synchronized manner.

3.Overall algorithm framework: Production status, deep reinforcement learning, and equipment execution loop

4. Core Competence 1: Integrated Perception of Production and Logistics Status

5. Core Competence 2: Dynamically generate the current transportation tasks that are worth performing

6. Core Competence 3: Intelligent Multi-Item Loading and Complete Assembly of Components

7. Core Competence 4: Cross-floor collaboration between AGVs and cargo elevators

8. How does the deep reinforcement learning scheduling algorithm complete a decision?

9. Deep reinforcement learning decision-making and dual-layer protection of industrial rules

10. Value to the management team: throughput, flexibility and scalability

III. Detailed Explanation of Content


1. The fifth-floor automotive parts workshop: Why logistics scheduling becomes a production bottleneck

In order to make the technical logic more in line with common discrete manufacturing scenarios, this paper constructs a five-layer automotive reducer component processing and assembly workshop. The main logistics objects include input shafts, intermediate shafts, gears, etc., which are batch parts. Different floors undertake functions such as blank preparation, rough machining, heat treatment, fine machining, and assembly. AGVs are responsible for intra-floor distribution, while freight elevators are responsible for vertical transportation between floors. As the order batch changes, the processing cycle and arrival rhythm of different parts are not completely consistent. Therefore, the logistics system needs to constantly make choices among "which type of part is lacking, which equipment can receive it, which vehicle is suitable for execution, and when is the freight elevator available".顺力智能物流

1.1  Five-floor demonstration of functional division


Floor

Main production functions

Typical logistics task

5F

Incoming materials, raw material preparation and raw material storage

Send the forged parts/rod materials/gear blanks that arrive in batches to the executable downstream processes.

4F

Initial processing: Turning, gear hobbing/grinding, drilling and milling, etc.

Distribute the materials among multiple parallel processing units and send the completed items to the heat treatment process.

3F

Heat treatment, cleaning, cooling and intermediate inspection

Coordinate the output of batch processes to avoid long-term material shortages or excessive accumulation in the subsequent fine processing stage.

2F

Grinding teeth, honing teeth, fine grinding and quality inspection

Send the qualified parts into the corresponding storage area to prepare all the necessary materials for the assembly of the first floor assembly.

1F

Assembly of complete sets, assembly of reducer assemblies, final inspection and warehouse dispatch

According to the assembly requirements, receive different combinations of parts and drive the assembly to be completed and rolled off the production line.

1.2 The real challenge is not "having vehicles available", but "the need for resources to be coordinated simultaneously"

There are multiple parallel equipment in the same process: if all vehicles tend to go to the nearby workstations, it is likely to cause local congestion, while other equipment is waiting.

The processing cycles of different parts are different: the completion of a large number of certain parts upstream does not mean that downstream assembly can continue; the matching relationship determines the truly scarce materials.

The freight elevator is a shared serial resource: when multiple floors simultaneously have cross-floor demands, the waiting, connection and service sequence will directly affect the transportation waiting time.

Multiple AGVs will compete for the same batch of inventory and the same target capacity: if only independent local decisions are made, it is easy to have repeated competing for tasks or resource conflicts.

Small batch and multi-variety production will constantly change "tasks": fixed priorities are difficult to adapt to the changes in order structure and work-in-process status in the long term.

Shunli Intelligent Warehouse Logistics Solution is precisely designed to address such multi-floor and multi-resource coupling problems.

2. Optimization objective: Not to make the equipment busier, but to ensure that more assemblies can be produced and released from the line in a synchronized manner.

In this scenario,the scheduling system is concerned with the number of reducer assemblies that are completed for final inspection and moved to the finished product area within a unit of time.This represents the throughput capacity of the entire production chain.Vehicle utilization rate,empty load rate,the number of cargo elevator movements,and workstation inventory levels can all be used as operational indicators,but they are not the goals.A seemingly"very busy"logistics system,if it consistently piles up a large number of parts on the third or second floor while the assembly area on the first floor is always lacking key components,the output will not increase.

From the management perspective,to judge whether the scheduling is good or not,one should not only look at"how many miles the AGV has traveled"or"how high the utilization rate of the cargo elevator is",but rather whether the logistics system is delivering the correct materials to the correct production processes at the right time,thereby reducing line waiting time and promoting product formation.

Level

Common local indicators

The issues that this algorithm focuses more on

AGV

Mileage, utilization rate, empty load rate

Does the vehicle operation transfer the current required materials to the next effective process?

Elevator

Number of runs, occupancy rate

Whether to provide effective services when real cross-layer demands arise and reduce unnecessary waiting time

Workstation

Inventory level, busy rate

Is there an imbalance between the upstream and downstream processes? Has there been a situation where one part accumulates while another part is short of materials?

Assembly

Completion rate, waiting time

Have all the key components arrived on schedule?

The entire line

The effective output of the reducer assembly within the given time period

顺力智能物流

Shunli Intelligent Bidirectional Concealed AGV Process Interchange Flow

3. Overall algorithm framework: Production status, deep reinforcement learning, and equipment execution loop

This algorithm is based on a prototype project of Shunli Intelligent. In the discrete manufacturing simulation software, a five-layer production and logistics environment is constructed, and a deep reinforcement learning scheduler is run outside the simulation environment. The simulation system continuously provides states such as vehicles, cargo elevators, workstations, inventory, and production progress. The algorithm generates transportation tasks based on the current state, and then undergoes verification through conditions such as process routes, inventory, capacity, and critical resource conflicts before being issued for execution. The new state after task completion re-enters the next round of judgment, forming a continuous loop.顺力智能物流

The closed loop from the production status to the deep reinforcement learning decision-making and then to the equipment execution

The key point of this structure is that the deep reinforcement learning algorithm is responsible for "finding better scheduling options in complex states", while industrial rules are responsible for "ensuring that actions can be executed by real devices". These two are not in a substitution relationship; instead, they are a combination of optimization capabilities and engineering constraints.

 4. Core Competence 1: Integrated perception of production and logistics status, enabling the algorithm to understand the workshop

From the perspective of deep reinforcement learning, the algorithm first needs to represent the current operation status of the workshop as a "state space". To facilitate a quick understanding of the core of the algorithm, the underlying feature encoding will not be elaborated here. The state space can be summarized as the position feature of the AGV, the task feature, the cargo and remaining capacity feature, the operation feature of the cargo elevator, the processing and caching status of the workstation, the material inventory and in-transit status, the assembly demand, etc. The algorithm continuously reads these information and determines which transportation actions are more likely to drive the entire production chain towards the final assembly line delivery.


State space category

Demonstration information

The question it answered

AGV

Location characteristics, task characteristics, cargo structure, remaining capacity, waiting status, etc.

Which vehicle is suitable for the mission? Does it have the conditions for continued transportation or consolidation?

Elevator

Floor location, operation status, waiting for elevator task, connection resource status, etc.

Is it currently suitable for cross-layer operation? Which tasks should the limited vertical transportation resources be prioritized to serve?

Workstation

Processing status, acceptable capacity, processing-in-progress and completed inventory, etc.

Can the target workstation receive the goods at present? Is there any local backlog?

Material flow

Inventory, in-transit quantity, downstream demand, waiting time, etc.

What else is lacking in the system at the moment? Which materials have exceeded the required quantity?

Assembly requirements

The availability of key components, assembly cycle time, and the batches awaiting assembly, etc.

Which transportation actions are more likely to directly lead to the completion of the assembly line?

From "having goods to move" to "valuable goods are worth moving" - this is also the case when there are 10 parts in the inventory. If the downstream of these parts has already accumulated, while another type of part is blocking the assembly completeness, the value of these two tasks for the overall line throughput is completely different.顺力智能物流

Shunli Intelligent Palletizing and Forking AGV Elevator for Material Handling

5. Core Competence 2: Dynamically generate the current transportation tasks that are worth performing

In multi-floor production,the transportation target is not a fixed location.Take the fourth floor rough processing as an example.The same part may be received by multiple parallel machines simultaneously;the set of tasks that an AGV can perform varies when it is empty,loaded,waiting for a stop,or just completing cross-floor transportation.The algorithm adopts the approach of"first forming the current executable candidates,and then selecting from the candidates by the deep reinforcement learning strategy",rather than randomly trying all the stations.Candidate generation takes into account factors such as process routes,inventory,target capacity,vehicle load,and the status of key resources:

First,filter out the floors and workstations that are not allowed to be reached based on the process route.

Then,filter out the current non-executable tasks based on the actual inventory,target capacity,vehicle load,and the occupation of key resources.

For the remaining candidate tasks,comprehensively evaluate the transportation distance,downstream material shortage,in-transit quantity,workstation load,and cross-floor costs.

After changes in production status,the candidate set automatically changes,so the strategy can reassign tasks in response to fluctuations in orders and work-in-progress.

What does this mean for the enterprise?When a production line expands parallel workstations,changes the order structure,or a workstation is temporarily busy,the scheduling logic does not have to rely entirely on manually rewriting a fixed set of station priorities,but can recompare the value of current tasks based on the real-time status.


6. Core Competence 3: Intelligent Multi-Item Loading and Complete Assembly of Components

In the production of automotive components, small batches and multiple varieties are often produced concurrently. When the input shaft, intermediate shaft, and different gears have only a small cross-layer demand respectively at the same production stage, if each type of material is transported separately by AGV and cargo lifts, it will repeatedly occupy the waiting space and vertical transportation opportunities. The algorithm allows combining multiple small batch demands into a single transportation task when the process direction compatibility, target direction consistency, and vehicle capacity are sufficient.顺力智能物流

6.1 "Stacking" is not simply stuffing the cargo to the brim

The combined materials must be compatible in the current transportation stage and the process route must not be disrupted to increase the loading rate.

The total vehicle load must meet the capacity limit,and there must also be sufficient receiving space at the target floor or workstation.

The scheduler should consider the actual downstream demand to avoid moving a large number of temporarily unused parts to the assembly area in advance,causing new backlogs.

After reaching the target floor,the materials can be further diverted according to the requirements of different workstations or complete sets areas,rather than requiring all materials to be sent to the same workstation.

In cross-floor scenarios,the effective combined transportation of value is not only a saving of AGV travel distance,but also has the opportunity to simultaneously reduce the repetitive processes of waiting for the elevator,entering and exiting the elevator,and cross-floor connection,making it more meaningful when vertical logistics resources are tight.

顺力智能物流

Shunli Intelligent Palletizing Picking AGV for Cross-Floor Transportation

7. Core Competence 4: Cross-floor collaboration between AGVs and cargo elevators

In the fifth-floor workshop, the cargo elevator is a shared vertical passage for all floors. If vehicles only select "near tasks" for transportation, they may push a large number of AGVs towards the elevator at the same time, causing congestion at the elevator waiting area; conversely, if the elevator only operates based on simple floor polling, it may frequently serve low-value tasks. The algorithm coordinates the cross-floor transportation intentions of AGVs with the current floor, queue, and production priority of the elevator in the same closed loop. This collaborative mechanism is the core component of the AGV flexible handling solution.

Typical situation

Problems that may arise from separate scheduling

The processing approach of collaborative scheduling

Multiple floors simultaneously request an elevator

The vehicles are queuing for a long time, and the key components are being delayed by routine tasks.

Based on the comprehensive ranking of production urgency, waiting queue sequence and cargo elevator position

The AGV only carries a small amount of materials and moves between floors.

The freight elevator is repeatedly occupied by small-scale tasks.

Prioritize the identification of loadable requirements and reduce the occurrence of duplicate cross-layer operations.

The downstream workstation is currently without capacity.

The vehicles were unable to unload the goods promptly after coming downstairs.

Include the target capacity in the feasibility judgment before crossing layers

装配区缺少关键配套件

The freight elevator is still serving the upstream regular replenishment tasks.

Increase the transportation priority directly related to the assembly completeness

顺力智能物流

Shunli Intelligent Palletizing and Retrieval AGV Automated Storage System for Warehouse Handling

8. How does the deep reinforcement learning algorithm make a scheduling decision?

顺力智能物流

The five business questions that the deep reinforcement learning scheduling algorithm continuously addresses

The core scheduling decision is implemented using the multi-agent deep reinforcement learning approach. To give a brief description, there is no need to understand the complex network structure. Just understand the algorithm output as five consecutive problems: who will transport, what to transport, where to deliver, how much to transport at a time, and when to cross floors. Corresponding to the algorithm, these problems constitute the main business meaning of the action space.

Question

Judgment criteria

The desired outcome

Who will handle it?

Vehicle position, idle/waiting status, cargo loading, task conflicts

Select the AGV that is currently more suitable for performing the task.

What to transport?

Inventory, in transit, downstream material shortage, complete assembly

Give priority to transporting the components that have a greater value in promoting production.

Where to deliver it?

Process route, parallel workstation status, target capacity

Dynamically select the currently more suitable processing or caching node

How many pieces will be transported at one time?

Vehicle remaining capacity, batch demand, target receiving capacity

Avoid excessive handling and make use of compatible container-sharing vehicles

When do you go across the floor?

Lift position, waiting line, urgency of task

The task of making the limited vertical transportation resources more valuable


8.1 The Basic Working Principle of Deep Reinforcement Learning

Deep reinforcement learning learns scheduling patterns by repeatedly interacting with the simulation environment: The algorithm reads the current state space, selects a set of scheduling actions, and the simulation environment executes them, returning new production and logistics states as well as training feedback. The training feedback is set around factors such as assembly output, effective production progress, waiting and congestion, and invalid actions, enabling the strategy to gradually learn which decisions are more conducive to improving the throughput per unit time.

8.2  Why adopt multi-agent deep reinforcement learning

Multiple AGVs and shared freight elevators will simultaneously change the system status:when one vehicle takes away the inventory,the tasks of other vehicles will change accordingly;when one vehicle enters the waiting area,it will also affect the subsequent cross-floor services.Multi-agent deep reinforcement learning enables multiple devices to learn to collaborate under the shared production goal,while retaining their ability to make quick decisions based on the current state.

The enterprise version can be understood as:during training,let the system learn to collaborate from"how the entire production line is operating";during operation,each device makes its own task selection based on the current state quickly.

9. Deep reinforcement learning decision-making with dual-layer guarantee of industrial rules

In industrial scenarios, deep reinforcement learning scheduling cannot merely focus on "algorithmic judgment being superior". Any task still must meet conditions such as process routes, actual inventory, vehicle capacity, target receiving capabilities, and critical resource exclusivity. Therefore, this algorithm separates the deep reinforcement learning decision-making from the industrial rule verification: the deep reinforcement learning algorithm is responsible for selecting more valuable tasks in complex states, while the execution layer is responsible for adhering to the unbreakable industrial boundaries.

顺力智能物流

Deep reinforcement learning decision-making and dual-layer guarantee mechanism based on industrial rules

Non-existent inventory cannot be"imagined",and materials that have been occupied by other tasks cannot be redistributed again.

Tasks such as AGV overload,target cache full load,and reverse process transportation will be directly filtered or rejected.

The connection area of the cargo elevator,waiting positions,and other key resources need to be used according to the mutually exclusive rules to avoid multiple vehicles competing simultaneously.

The actual results after task execution will be re-fed back to the scheduler,and the next round of decision-making will be based on the new state.

Tags