- Time: 2020 Feb 24 (Monday)
- Location: Location: Atkinson Auditorium @ UCSD
- Live stream recording available on youtube. (almost 8 hours!)
- Monocular 3D learning is hot and differentiable rendering is a useful tool for it (now available in pytorch3D)
- Learning on video (efficient training, 4K bird, oops! dataset)
- Self supervised learning for pretraining
- More work on understanding human intention
- And some solid work on incorporating conventional knowledge into CNN (charality net, GAN paint, make CNN shift invariant again, etc)
- Many female speakers. I noticed that the female researchers tend to work on projects that has more direct social impact (detecting birds, solving puzzles for archeology, migrate human dense pose for Chimpanzees, anonymize videos). To be fair, there are male researchers working on bias in DL too!
- 8:00 - 8:12 Justin Johnson End-to-End View Synthesis from a Single Image
- Feifei Li's student, UMich
- SynSin: End-to-end View Synthesis from a Single Image
- Input:
- one image
- transform based on given translation and rotation
- Challenge: need to know the depth, impainting missing regions
- Train without depth GT
- Hallucinating the occluded part
- Differentiable projection, Z-buffer --> PyTorch3D
- Similar to depth prediction (SfMLearner)?
- Next step: self-supervised depth

- 8:12 - 8:24 Alex Schwing Chirality Nets for Human Pose Regression
- UIUC
- Chirality Nets for Human Pose Regression NIPS 2019
- seeing the unseen (occlusion), and predict the future
- Pose regression tasks:
- concatenation of joint coordinates
- 2D pose mirrored --> 3D pose mirrored (charality equivariance)
- Work the charality equivariance into the neural nets
- Benefits of charality
- Data efficient
- Equivariance guarantee
- Similar to pointNet
- Will take you down for asymmetric tasks
- How about adding the self consistency loss in training?

- 8:24 - 8:36 Georgia Gkioxari Challenges in PyTorch 3D
- FAIR
- 2D to 3D: Mesh R-CNN
- Mask RCNN: more than 500 samples per image
- Batching in 3D: varying number of vertexes, hetereogeneous batches
- Auxiliary tensor to remember index
- Padded tensor
- different batching mode even int he same forward pass
- Differentiable renderer
- PyTorch3D is important for any Neural nets with rendering
- can be useful for 3D MOD
- video demo and another one

- 8:36 - 8:48 Carl Vondrick Oops! Predicting Unintentional Action
- Oops! Predicting Unintentional Action in Video CVPR 2020
- Survivor bias of video Data
- Opps! Dataset: failed videos
- What are the perception clues to tell intentionality
- Self-supervised clue: order, video speed
- speed of action alters perceptual judgement
- Video: intentional, transition, unintentional
- 8:48 - 9:00 Hyun Soo Park HUMBI dataset multiview human expression
- Univ of Minnesota
- Multiview data from diverse identities
- Collect garment, gaze, guesture
- $1 reward for participants
- 9:00 - 9:12 Rita Cucchiara Self-attention in Vision and Language
- 9:12 - 9:24 Vineeth N Balasubramanian Zero-shot Task Transfer/Causal Attributions in Neural Networks
- IIT Hyderabad
- related to and inspired by Taskconomy
- Input: task correlation matrix: either crowd sourcing or Taskconomy
- Zero shot: how to learn models for tasks without Gt, given other tasks with GT
- Meta Learning to learn a regressor to learn the weights for a new task based on the weight of existing tasks
- Limitation: output space need to be the same

- 9:24 - 9:36 Roozbeh Mottaghi Interactive Scene Understanding
- Allen Institute for AI
- CV problems: passive/active/interactive
- passive: fixed dataset
- active: change dataset to corresponds to the label
- interactive: no fixed dataset. The dataset depends on your actions. navigation, interactive QA
- Dataset: AI2Thor, AI2Thor-robo (sim2real dataset, inverted virtual KITTI)
- 9:36 - 9:48 Philipp Kraehenbuehl Objects as Points
- Efficiently Training Video Models A Multigrid Method for Efficiently Training Video Models
- Learning by Cheating
- Tracking objects as points, extremely fast on MOT

- 9:48 - 10:00 David Fouhey Internet-Scale Hands In Interaction
- 10:30 - 10:42 Laura Leal-Taixé Video anonymization
- 10:42 - 10:54 Vinay Namboodiri Integrating Vision, Speech and Language
- IIT Kampur
- Modify sync to improve lip sync
- Modify video to improve lip sync in different languages
- 10:54 - 11:06 Adriana Kovashka Reasoning about Complex Media from Weak Multi-Modal Supervision
- University of Pittsburgh
- Understanding of advertisement
- 11:06 - 11:18 Amir Zamir Perception in the Action Loop
- student of Silvio Savarese and Jitendra Malik.
- Datasets: fixed, passive
- Perception for active agents
- Having vidual priors helps visual navigation
- 11:18 - 11:30 Ayellet Tal Solving Jigsaw Puzzles
- Technion
- Archeologists are solving puzzles
- Reconstruct an object from a set of non-overlapping and unordered parts (NP-hard)
- Solving puzzles: take in image patches, then generate a coherent image
- Boundaries are eroded (for archeology use)
- Key ideas:
- Inpaint
- Learn to classify a pair is neighbor or not
- Even humans would need to know the maximum margin of erosion to perform well.

- 11:30 - 11:42 Boqing Gong Long-Tailed Visual Recognition Is A Domain Adaptation Problem
- 11:42 - 11:54 Chen Sun Speech2Action
- Annotating actions in videos are very costly
- Use human speech as weak supervision
- 1:00 - 1:12 Andreras Geiger Learning Implicit 3D Reconstruction without 3D Supervision
- Toyota Institute, University of Tubingen, MPI
- Output representation:
- Voxels/Points/Meshes (discretize)
- Occupancy networks
- Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision
- No need to store intermediate results
- blog

- 1:12 - 1:24 Hamed Pirsiavash Adversarial Patches Exploiting Contextual Reasoning in Object Detection
- Univ of Maryland
- last layer has receptive field of the entire image. YOLO uses global info
- conditional reasoning leaves backdoor for adversarial attacks
- single shot detectors are hard to defend against adversarial patches
- blinding toward one or multiple object classes
- Introduce DetGrad-Cam
- Impossible to detect the patch (the patch can be initialized from a natural image and it does not need to be very random)

- 1:24 - 1:36 Hao Su SAPIEN: A Simulated Part-based Interactive Environment
- Simulation environment
- 1:36 - 1:48 Jun-Yan Zhu Visualizing and Understanding GANs
- Seeing What a GAN Cannot Generate ICCV 2019
- which neurals are responsible for a feature
- How to manipulate an existing photo? Reconstruct the image first then manipulate the neurons

- 1:48 - 2:00 Manmohan Chandraker Physically-Based Learning for Inverse Rendering
- 2:00 - 2:12 Minsu Cho Composing Neural Features for Visual Correspondence in the Wild
- 2:12 - 2:24 Natalia Neverova Transferring Dense Pose to Animals
- FAIR in Paris, MPI
- DensePose (ultimate human parsing)
- bijective mapping between 3D models of human and other creatures
- Temporal consistency is a good indicator of network performance as well
- 2:24 - 2:36 Octavia Camps Compact and Interpretable Dynamics-based Video Representations
- Integrate kalman filter into the representation and keep the representations consistent
- 2:36 - 2:48 Rei Kawakami Improving Robustness in Recognition: Motion, Open-set, and Multi-task learning 川上 玲(東京大学)
- Video can help object detection as well -- the motion pattern can help differentiate FPs
- みなさんのデータによる首里城のデジタル復元 (inspired by building Rome in a day)
- 2:48 - 3:00 Vicente Ordonez Explicit Compositionality in Language and Vision
- 3:30 - 3:42 Richard Zhang CNN-Generated Images Are Surprisingly Easy to Spot ... for Now
- 3:42 - 3:54 Sanjeev Koppal Fast Foveating Cameras, LIDARs and Projectors
- 3:54 - 4:06 Shuran Song Grasping in the Wild












