1. X
  2. Punyajoy Saha
Log inOpen app
Punyajoy Saha
426 posts
user avatar
Punyajoy Saha
@punyajoysaha
Cheif Engineer, Samsung Research Institute Bangalore | PhD@CNeRG |Ex-Intern CLAWS group @GeorgiaTech ,LT group @unihh . NLP | Safety | Social Good.
Kharagpur, India
punyajoy.github.io
Joined December 2016
612
Following
240
Followers
RepliesRepliesMediaMedia

Get the full app experience

Unlock more features and see what people are talking about right now.

Open X
  • Pinned
    user avatar
    Punyajoy Saha
    @punyajoysaha
    Mar 10
    🚨 New preprint: Can Safety Emerge from Weak Supervision? We study whether small LLMs can learn safer behavior without large human-annotated datasets. We introduce Self-MOA, an automated alignment framework. 📄 arxiv.org/pdf/2603.07017 Thread 👇
    1
    1
    91
  • user avatar
    Punyajoy Saha
    @punyajoysaha
    Oct 5, 2024
    The dataset, code and the preprint are available now. Feel free to use the data in your work. Preprint :- arxiv.org/abs/2410.01400 Code :- github.com/hate-alert/Cro…
    user avatar
    Punyajoy Saha
    @punyajoysaha
    Sep 25, 2024
    🚨 New Dataset Alert: CrowdCounter! 🚨 A groundbreaking dataset with 3,425 hate speech-counterspeech pairs across 6 styles (empathy, humor, etc.). Tested top models—Flan-T5 excelled at standard responses, DialoGPT for type-specific. 💬🙌 #AI #Counterspeech #TechForGood
    1
    149
  • user avatar
    Punyajoy Saha
    @punyajoysaha
    Sep 25, 2024
    🚨 New Dataset Alert: CrowdCounter! 🚨 A groundbreaking dataset with 3,425 hate speech-counterspeech pairs across 6 styles (empathy, humor, etc.). Tested top models—Flan-T5 excelled at standard responses, DialoGPT for type-specific. 💬🙌 #AI #Counterspeech #TechForGood
    user avatar
    CNeRG IIT KGP
    @cnerg
    Sep 25, 2024
    Replying to @cnerg
    CrowdCounter: A benchmark type-specific multi-target counterspeech dataset @punyajoysaha, Abhilash Datta, Abhik Jana, @Animesh43061078 3/3
    1
    3
    490
  • user avatar
    Punyajoy Saha
    @punyajoysaha
    May 22, 2024
    I am at @LrecColing 2024, presenting two works based on LLMs for content moderation on 23rd May Poster Area II, 9-10:40 CEST. Feel free to reach out.. 1. "On Zero-Shot Counterspeech Generation by Large Language Models" - explores LLMs for generating counterspeech.
    1
    7
    316
  • user avatar
    Punyajoy Saha
    @punyajoysaha
    Mar 27, 2024
    We are doing some exciting work in understanding how we can make the multimodal hate speech detection explainable ... Please apply if you are interested. You can DM me if you have any questions ...
    user avatar
    Animesh Mukherjee
    @Animesh43061078
    Mar 27, 2024
    One JRF post available for research on explainable hate speech detection models from multimodal content. Application deadline: 4/4/2024. Please apply through the link: erp.iitkgp.ac.in/SRICStaffRecru… @punyajoysaha @iMithun_Das @NaqueeRizwan @cnerg #NLProc
    1
    1
    313