Related papers: The Advantage of a Multi-User Mode
Multi-person motion prediction is a challenging task, especially for real-world scenarios of highly interacted persons. Most previous works have been devoted to studying the case of weak interactions (e.g., walking together), in which…
A visualization scheme for quantum many-body wavefunctions is described, which we have termed qubism. Its main property is its recursivity: increasing the number of qubits reflects in an increase in the image resolution. Thus, the plots are…
Photons are promising candidates for quantum information technology due to their high robustness and long coherence time at room temperature. Inspired by the prosperous development of photonic computing techniques, recent research has…
Multi-mode expansions in computational quantum dynamics promise convergence toward exact results upon increasing the number of modes. Convergence is difficult to ascertain in practice due to the unfavourable scaling of required resources…
Existing approaches for the design of interpretable agent behavior consider different measures of interpretability in isolation. In this paper we posit that, in the design and deployment of human-aware agents in the real world, notions of…
The present work investigates whether different quantification mechanisms (set comparison, vague quantification, and proportional estimation) can be jointly learned from visual scenes by a multi-task computational model. The motivation is…
Today millions of mobile apps are downloaded and used all over the world. Guidelines and best practices on how to design and develop mobile apps are being periodically released, mainly by mobile platform vendors and researchers. They cover…
Multimodal sentiment analysis is an important area for understanding the user's internal states. Deep learning methods were effective, but the problem of poor interpretability has gradually gained attention. Previous works have attempted to…
A good evaluation framework should evaluate multimodal machine translation (MMT) models by measuring 1) their use of visual information to aid in the translation task and 2) their ability to translate complex sentences such as done for…
Single photon emitters often rely on a strong nonlinearity to make the behaviour of a quantum mode susceptible to a change in the number of quanta between one and two. In most systems the strength of nonlinearity is weak, such that changes…
Explainability is motivated by the lack of transparency of black-box Machine Learning approaches, which do not foster trust and acceptance of Machine Learning algorithms. This also happens in the Predictive Process Monitoring field, where…
The widely accepted basis for quantum computing advantage is derived from the entanglement and superposition properties of the probabilistic interpretation of the underlying quantum mechanical formalism which in turn is widely accepted…
Evidence for fine-tuning of physical parameters suitable for life can perhaps be explained by almost any combination of providence, coincidence or multiverse. A multiverse usually includes parts unobservable to us, but if the theory for it…
Vision based human motion recognition has fascinated many researchers due to its critical challenges and a variety of applications. The applications range from simple gesture recognition to complicated behaviour understanding in…
In optical interferometry multi-mode entanglement is often assumed to be the driving force behind quantum enhanced measurements. Recent work has shown this assumption to be false: single mode quantum states perform just as well as their…
Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that take historical frames and current…
Multimodal interfaces, combining the use of speech, graphics, gestures, and facial expressions in input and output, promise to provide new possibilities to deal with information in more effective and efficient ways, supporting for instance:…
Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text,…
In this work, we aim to improve the 3D reasoning ability of Transformers in multi-view 3D human pose estimation. Recent works have focused on end-to-end learning-based transformer designs, which struggle to resolve geometric information…
Hamiltonian of a parametric process describing the interaction of a finite number of optical cavity modes, with the microwave field is proposed. Three-boson interaction is due to electro-optical effect. Analysis of the model is based on the…