Video Experimental Relacionado
Updated: Jan 17, 2026

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
1.0K
Refinamiento Reforzado con Expansión Autoconsciente para Conducción Autónoma de Extremo a Extremo
IEEE transactions on pattern analysis and machine intelligence
|January 14, 2026
Resumen
El Refinamiento Reforzado con Expansión Autoconsciente (R2SE) mejora la conducción autónoma de extremo a extremo al refinar escenarios desafiantes y al mismo tiempo retener políticas de conducción generales. Este enfoque mejora la seguridad y la robustez de los sistemas de autoconducción.
Área de la Ciencia:
- Inteligencia Artificial
- Robótica
- Ciencias de la Computación
Sus antecedentes:
- Los modelos de conducción autónoma de extremo a extremo mapean datos de sensores a acciones de conducción.
- Los modelos existentes de aprendizaje por imitación (IL) tienen dificultades para generalizar a situaciones de conducción difíciles y carecen de retroalimentación posterior al despliegue.
- El aprendizaje por refuerzo (RL) puede abordar casos complejos, pero a menudo conduce a sobreajuste y al olvido del conocimiento general.
Objetivo del estudio:
- Introducir un nuevo pipeline de aprendizaje, Refinamiento Reforzado con Expansión Autoconsciente (R2SE), para la conducción autónoma de extremo a extremo.
- Mejorar la generalización a casos difíciles y garantizar la mejora continua de las políticas de conducción.
- Superar las limitaciones de los enfoques actuales de IL y RL en la conducción autónoma.
Principales métodos:
- R2SE emplea un pipeline de tres componentes: Preentrenamiento Generalista con asignación de casos difíciles, Ajuste Fino Especialista Reforzado Residual y Expansión Adaptadora Autoconsciente.
- El Preentrenamiento Generalista identifica los casos propensos a fallos para un refinamiento específico.
- El Ajuste Fino Especialista Reforzado Residual utiliza RL para optimizar el rendimiento en dominios difíciles, preservando al mismo tiempo el conocimiento general.
Principales resultados:
- R2SE demuestra una generalización, seguridad y robustez de políticas de largo alcance mejoradas en comparación con los sistemas de extremo a extremo (E2E) de última generación.
- El método refina eficazmente el rendimiento en escenarios de conducción desafiantes.
- Los resultados experimentales se validaron tanto en simulaciones de bucle cerrado como en conjuntos de datos del mundo real.
Conclusiones:
- El refinamiento reforzado ofrece una estrategia eficaz para sistemas de conducción autónoma escalables.
- R2SE permite la mejora continua de las políticas de conducción mediante la integración dinámica de conocimientos especializados.
- El pipeline propuesto aborda los desafíos clave en la generalización y la robustez para la conducción autónoma E2E.
Palabras clave:
conducción autónomaaprendizaje por refuerzoaprendizaje por imitaciónredes neuronales profundassistemas de controlMás Videos Relacionados
Videos de Conceptos Relacionados
Introspection
208
Introspection, long upheld as a reliable route to self-knowledge, involves examining one's thoughts, emotions, and mental processes. It underpins many psychological practices, from mindfulness meditation to psychotherapy and self-help strategies. However, empirical evidence challenges the accuracy of introspection as a means of understanding oneself.Limitations of Introspective InsightSeminal work by Nisbett and Wilson demonstrated that individuals are frequently unaware of the true causes...
208
Automatic Processing and Automatic Social Behavior
215
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
215
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
3.8K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.8K
Elaborative Rehearsals
338
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
The effectiveness of...
338
Reinforcement
839
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
839
Controller Configurations
352
Controller configurations are crucial in a car's cruise control system because they manage speed over time to maintain a consistent pace regardless of road conditions, thereby meeting design goals. In traditional control systems, fixed-configuration design involves predetermined controller placement. System performance modifications are known as compensation.
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
352

