DEEP REINFORCEMENT LEARNING CONTROL OF A TENSION-CONSTRAINED TENDON-DRIVEN UNDERACTUATED ROBOTIC FINGER: A COMPARATIVE STUDY OF DDPG AND SAC

Authors

DOI:

https://doi.org/10.37943/CVPA1758

Keywords:

prosthetic finger; reinforcement learning; underactuated control; tendon-driven systems; actuation constraints; soft actor–critic

Abstract

Tendon-driven robotic fingers are underactuated and strongly nonlinear, and their unidirectional, friction-attenuated tendon tension imposes actuation constraints that any practical controller must respect. A previous study established that a deterministic learning controller (DDPG) can track a grasp trajectory for such a finger and compared it against model-based schemes. The present work extends that line of investigation by asking a narrower but unresolved question: for this same single-actuator, three-link tendon-driven architecture, how does a deterministic actor-critic policy (DDPG) compare with a stochastic, maximum-entropy policy (SAC) once a tension-limit penalty is built directly into the learning objective? Using the previously validated physics-based simulation model, both agents are trained under identical conditions with a reward that jointly penalizes tracking error and tension-constraint violation, and are evaluated across random, sinusoidal, and step disturbances using steady-state error, RMS, ISE, and IAE. The results show a consistent and interpretable trade-off: the stochastic SAC policy attains lower tracking error and better disturbance rejection, especially under time-varying and abrupt loads, at the cost of noisier control and longer training, whereas the deterministic DDPG policy converges faster and produces smoother torque and tension profiles but adapts less well to coupled disturbances. The study isolates the role of the exploration mechanism and the tension penalty in shaping these behaviors and offers practical guidance on selecting between deterministic and stochastic actor-critic methods for physically constrained, underactuated tendon-driven systems.

References

Zhang, T., Zheng, K., Tao, H., & Liu, J. (2025). A Soft Wearable Modular Assistive Glove Based on Novel Miniature Foldable Pouch Motor Unit. Advanced Intelligent Systems, 7(11), 2500274. https://doi.org/10.1002/aisy.202500274

Palli, G. (2007). Model and control of tendon actuated robots.

Borghesan, G., Palli, G., & Melchiorri, C. (2010, May). Design of tendon-driven robotic fingers: Modeling and control issues. In 2010 IEEE International conference on robotics and automation (pp. 793-798). IEEE. https://doi-org.ezproxy.utlib.ut.ee/10.1109/ROBOT.2010.5509899

Suleimenov, K., Kapsalyamov, A., Abdikenov, B., Ozhikenova, A., Igembay, Y., & Ozhikenov, K. (2025). Comparative Analysis of Model-Based and Data-Driven Control for Tendon-Driven Robotic Fingers. Mathematics, 13(22), 3669. https://doi.org/10.3390/math13223669

Astrom, K. J. (1995). PID controllers: theory, design, and tuning. The international society of measurement and control. https://www.wiley.com/en-us/shop/general-introductory-industrial-engineering/pid-controllers-theory-design-and-tuning-2nd-edition-p-9781556175169

Isidori, A., van Schuppen, J., Sontag, E., Thoma, M., & Krstic, M. (1995). Communications and control engineering. Nonlinear control systems. https://doi.org/10.1007/978-1-84628-615-5

Alam, U. K., Shedd, K., & Haghshenas-Jaryani, M. (2023). Trajectory control in discrete-time nonlinear coupling dynamics of a soft exo-digit and a human finger using input–output feedback linearization. Automation, 4(2), 164-190. https://doi.org/10.3390/automation4020011

Capotondi, M., Turrisi, G., Gaz, C., Modugno, V., Oriolo, G., & De Luca, A. (2020). Learning feedback linearization control without torque measurements. In 2020 I-RIM Conference. https://doi.org/10.5281/zenodo.4781489

Málik, R., Okienková, K., Vasko, D., Kardoš, J., Tárník, M., & Paulusová, J. (2024). Robust Control of a MIMO Laboratory Motion System Using Variable Structure Control and the Computed Torque Method With the Velocity Limitation. IEEE Access, 12, 143869-143882. https://doi.org/10.1109/ACCESS.2024.3471678

. Simonini, G., Baracca, M., Cavaliere, T. V., Bicchi, A., & Salaris, P. (2025). A novel formulation for adaptive computed torque control enabling low feedback gains in highly dynamical tasks. IEEE Access.https://doi.org/10.1109/ACCESS.2025.3561635

Utkin, V. I. (2013). Sliding modes in control and optimization. Springer Science & Business Media.

Van, M., & Ge, S. S. (2020). Adaptive fuzzy integral sliding-mode control for robust fault-tolerant control of robot manipulators with disturbance observer. IEEE Transactions on Fuzzy Systems, 29(5), 1284-1296. https://doi.org/10.1109/TFUZZ.2020.2973955

Shtessel, Y., Edwards, C., Fridman, L., & Levant, A. (2014). Sliding mode control and observation (Vol. 10, pp. 978-970). New York: Birkhäuser. https://doi.org/10.1007/978-0-8176-4893-0

Dirara, H. G., Yareshe, F. T., & Abdissa, C. M. (2025). Design and analysis of adaptive fuzzy super-twisting sliding mode controller for uncertain 2-DOF robotic manipulator. IEEE Access. https://doi.org/10.1109/ACCESS.2025.3581449

Lochan, K., Seneviratne, L., & Hussain, I. (2025). Adaptive global super-twisting sliding mode control for trajectory tracking of two-link flexible manipulators. IEEE Access. https://doi.org/10.1109/ACCESS.2025.3557202

Li, L., Su, Y., Kong, L., Jiang, K., & Zhou, Y. (2022). TDE-based adaptive super-twisting multivariable fast terminal slide mode control for cable-driven manipulators with safety constraint of error. IEEE Access, 11, 6656-6664. https://doi.org/10.1109/ACCESS.2022.3232555

Wang, Y., Yan, F., Chen, J., & Chen, B. (2018). Continuous nonsingular fast terminal sliding mode control of cable-driven manipulators with super-twisting algorithm. IEEE Access, 6, 49626-49636. https://doi.org/10.1109/ACCESS.2018.2868988

Sutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction (Vol. 1, No. 1, pp. 9-11). Cambridge: MIT press. https://mitpress.mit.edu/9780262193986/reinforcement-learning/

Sumiea, E. H., Abdulkadir, S. J., Alhussian, H. S., Al-Selwi, S. M., Alqushaibi, A., Ragab, M. G., & Fati, S. M. (2024). Deep deterministic policy gradient algorithm: A systematic review. Heliyon, 10(9). https://doi.org/10.1016/j.heliyon.2024.e30697

Elumalai, V. K. (2025). A proximal policy optimization based deep reinforcement learning framework for tracking control of a flexible robotic manipulator. Results in Engineering, 25, 104178. https://doi.org/10.1016/j.rineng.2025.104178

Viswanadhapalli, J. K., Elumalai, V. K., Shah, S., & Mahajan, D. (2024). Deep reinforcement learning with reward shaping for tracking control and vibration suppression of flexible link manipulator. Applied Soft Computing, 152, 110756. https://doi.org/10.1016/j.asoc.2023.110756

Rúbio, G. D. P., Costa, M. C. B., & Vimieiro, C. B. S. (2025). A New Proposal for Intelligent Continuous Controller of Robotic Finger Prostheses Using Deep Deterministic Policy Gradient Algorithm Through Simulated Assessments. Robotics, 14(4), 49. https://doi.org/10.3390/robotics14040049

. Kapsalyamov, A., Brown, N. A., Goecke, R., Jamwal, P. K., & Hussain, S. (2025). Velocity control of a Stephenson III six-bar linkage-based gait rehabilitation robot using deep reinforcement learning. Neural computing and applications, 37(7), 5671-5682. https://doi.org/10.1007/s00521-024-10944-2

Wu, J., Chen, S., Wang, Y., Zhang, P., Su, C. Y., & Sato, D. (2024). A transfer reinforcement learning-based control method for vertical n-link underactuated manipulator with two passive joints. IEEE/ASME Transactions on Mechatronics, 30(3), 1797-1806. https://doi-org.ezproxy.utlib.ut.ee/10.1109/TMECH.2024.3419803

He, W., Gao, H., Zhou, C., Yang, C., & Li, Z. (2020). Reinforcement learning control of a flexible two-link manipulator: An experimental investigation. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 51(12), 7326-7336. https://doi-org.ezproxy.utlib.ut.ee/10.1109/TSMC.2020.2975232

Fujimoto, S., Hoof, H., & Meger, D. (2018, July). Addressing function approximation error in actor-critic methods. In International conference on machine learning (pp. 1587-1596). PMLR. https://doi.org/10.48550/arXiv.1802.09477

Aydogmus, O., & Yilmaz, M. (2023). Comparative analysis of reinforcement learning algorithms for bipedal robot locomotion. IEEE Access, 12, 7490-7499. https://doi.org/10.1109/ACCESS.2023.3344393

Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018, April). Deep reinforcement learning that matters. In Proceedings of the AAAI conference on artificial intelligence (Vol. 32, No. 1). https://doi.org/10.1609/aaai.v32i1.11694

Ziebart, B. D. (2010). Modeling purposeful adaptive behavior with the principle of maximum causal entropy. Carnegie Mellon University.

Zollo, L., Roccella, S., Guglielmelli, E., Carrozza, M. C., & Dario, P. (2007). Biomechatronic design and control of an anthropomorphic artificial hand for prosthetic and robotic applications. IEEE/ASME Transactions On Mechatronics, 12(4), 418-429. https://doi.org/10.1109/TMECH.2007.901936

Nwachukwu, S. E., Folly, K. A., & Awodele, K. O. (2025). A comparative study between soft actor-critic (SAC) and deep deterministic policy gradient (DDPG) algorithms for solar PV MPPT control under partial shading conditions. IEEE Access. https://doi.org/10.1109/ACCESS.2025.3561807

Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018, July). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning (pp. 1861-1870). Pmlr. https://doi.org/10.48550/arXiv.1801.01290

Rusinak, D., Kelemen, M., & Virgala, I. (2026). A Systematic Benchmark of Reinforcement Learning Formulations for Tendon-Driven Continuum Robot Control. IEEE Access. https://doi.org/10.1109/ACCESS.2026.3690990

Downloads

Published

2026-09-30

How to Cite

Suleimenov, K., Beisembekova , R., Abdikenov, B., Ozhikenova, A. ., Ozhiken, A., & Kapsalyamov, A. . (2026). DEEP REINFORCEMENT LEARNING CONTROL OF A TENSION-CONSTRAINED TENDON-DRIVEN UNDERACTUATED ROBOTIC FINGER: A COMPARATIVE STUDY OF DDPG AND SAC. Scientific Journal of Astana IT University, 27(3), 138–151. https://doi.org/10.37943/CVPA1758

Issue

Section

Information Technologies